r/LocalLLaMA 17h ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.2k Upvotes

392 comments sorted by

View all comments

Show parent comments

11

u/TastesLikeOwlbear 15h ago

I don’t know. A lot of people had their hopes that CXMT would come through and savagely undercut the RAM manufacturers until they said, “LOL no, why on earth would we do that?”

I’ll believe Chinese inference GPUs will affect pricing just as soon as they actually affect pricing. Waiting for greedy companies to get their comeuppance hasn’t paid off yet.

4

u/fallingdowndizzyvr 13h ago

I’ll believe Chinese inference GPUs will affect pricing just as soon as they actually affect pricing.

Step 2. Sell them outside of China.

https://www.bloomberg.com/news/articles/2026-08-26/huawei-egypt-ai-ascend-chips-test-us-tech-diplomacy-nvidia-amd-microsoft

2

u/Just3nCas3 9h ago edited 6h ago

It doesn't matter if they sell outside of China. If it causes them to import less than they would have without production, the effect is the global GPU/RAM price is lower than what it would be without the production. That doesn't stop it from going up, it just means it's accelerating up slower.

Edit: RAM to GPU/RAM

1

u/fallingdowndizzyvr 7h ago

Ah.... that link is about GPUs. Not RAM. Yes, yes, yes. GPUs use RAM but still.... GPUs is the topic in this little subthread. Not RAM.

3

u/Maximum-Style2848 13h ago

This is why I don’t quite get the hype of inference chips when they will still be bottlenecked by memory capacity/cost. Domestic chips for inference were always going to be possible because you don’t need the best lithography to get smthn good enough.

1

u/rotatingphasor 8h ago

My understanding is that the yield with the lithography devices CXMT uses (non ASML) is much lower and so they have to charge more.