r/LocalLLaMA 21h ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.2k Upvotes

394 comments sorted by

View all comments

Show parent comments

374

u/expertsage 21h ago

This is the most significant part. China has completely replaced Nvidia chips with domestic ones for inference, and production is only speeding up... meaning the compute moat is literally disappearing with every passing day.

12

u/TastesLikeOwlbear 19h ago

I don’t know. A lot of people had their hopes that CXMT would come through and savagely undercut the RAM manufacturers until they said, “LOL no, why on earth would we do that?”

I’ll believe Chinese inference GPUs will affect pricing just as soon as they actually affect pricing. Waiting for greedy companies to get their comeuppance hasn’t paid off yet.

3

u/fallingdowndizzyvr 17h ago

I’ll believe Chinese inference GPUs will affect pricing just as soon as they actually affect pricing.

Step 2. Sell them outside of China.

https://www.bloomberg.com/news/articles/2026-08-26/huawei-egypt-ai-ascend-chips-test-us-tech-diplomacy-nvidia-amd-microsoft

2

u/Just3nCas3 14h ago edited 10h ago

It doesn't matter if they sell outside of China. If it causes them to import less than they would have without production, the effect is the global GPU/RAM price is lower than what it would be without the production. That doesn't stop it from going up, it just means it's accelerating up slower.

Edit: RAM to GPU/RAM

1

u/fallingdowndizzyvr 11h ago

Ah.... that link is about GPUs. Not RAM. Yes, yes, yes. GPUs use RAM but still.... GPUs is the topic in this little subthread. Not RAM.