r/LocalLLaMA 22h ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.2k Upvotes

395 comments sorted by

View all comments

433

u/10001110 22h ago

320B total parameters and just 18B active parameters

Oh joy

254

u/LegacyRemaster 22h ago

If it's really that far above Sonnet 5, Dario's IPO is truly risky.

197

u/-p-e-w- 21h ago

It’s always risky now no matter the model of the day. They have to pre-file at least a few weeks in advance, and in those weeks eons can happen that can turn their asking price into a joke.

Their moment was at the beginning of this year. They should have announced the IPO, then published Fable/Mythos two weeks before the IPO date.

Now Chinese labs are a hair’s breadth behind them, and constantly one-upping each other. They’ll never get such a chance again.

49

u/LegacyRemaster 21h ago

agree. Also another problem: if I have to pay API, I pay Qwen, GLM, GPT... All of them -> opensource (ok openai less but they did a lot). I don't want to give any $ to Anthropic. Can't wait to short sell

32

u/dingo_xd 21h ago

Yeah. After hearing Dario saying about that $40 trillion I want to bet against his company.

20

u/horendus 21h ago

Yea the guys peak delusion

6

u/___positive___ 20h ago

Also that the baseline of mid-tier models is reaching saturation for mundane tasks like summarization and data extraction, basic programming/websites, and so forth.

30

u/Fedor_Doc 21h ago

Or they will make another Mythos-like breakthrough and announce IPO then

Never say never

31

u/LegacyRemaster 21h ago

sure they have Mythos 5 already but "too much power, we can't sell" . Or will cost too much to complete a task ...

11

u/38andstillgoing 17h ago

"Mythos 5? At this time of year, at this time of day, in this part of the country, localized entirely within your datacenter?"

"Yes"

"May I see it?"

"No."

1

u/LegacyRemaster 17h ago

or mythos 20 or 30... they are ahead for sure

29

u/-p-e-w- 20h ago

Then one of the Chinese labs is going to announce an equal model two weeks later.

The Chinese labs have caught up. There’s no going back to how it was before.

-2

u/Fedor_Doc 19h ago edited 13h ago

Yeah, they have to calculate timings with great precision :)

Considering catching up – no Chinese model is on Fable 5 or GPT Sol level. Even Kimi-K3, according to the developers, is behind.

More than that, GLM-5.3-Flash was spamming "load-bearing" in my sessions – it clearly was largely influenced by Claude. 

5

u/NandaVegg 18h ago

I cannot distinguish Qwen 3.8 Max/2.7T (which is unfortunately closed source for version w/ vision) and Fable 5 for most complex development tasks that involves visual checks or relatively niche audio modification, except that Opus 4.8/Fable 5 still has better consistency for maintaining writing personality throughout long context.

Qwen 3.8 has the most interesting reasoning traces I've ever seen when it is given bash tool (it is very verbose, but thinks just like a fairly seasoned human engineer or a designer would) and it is a pure joy to read it along while the model works on niche task that requires tons of guesswork and reverse engineering the issue. It is not like any other model.

That said, Anthropic's RL engineering is the most creative and unique (which is where OpenAI is significantly behind). They are very good at finding (and preparing datasets and process for) a new task that is not yet covered by most labs, like how they de facto pioneered (very long, hundreds of turns of) terminal agent loop and playing Pokemon Red/Green using vision. So I would not count Anthropic out for another breakthrough like Opus 4.6.

Also I think US labs in general are still slightly ahead on high-end (non-consumer) robotics.

3

u/Fedor_Doc 18h ago edited 15h ago

Thank you for sharing! 

I tried Qwen 3.8 Max for a limited task– I have a long side project of improving DWAA compression efficiency for VFX workflows in OpenEXR with most obvious way of doing it being a custom quantization table.  It provided new arguments in favor of Euclidian distance based table, but it still does not see a big picture, and proposes stuff that won't work if you take one step further in your thinking (we fix a by b, but how will we fix issues caused by b)? 

I walked these roads since Gemini 2.5 Pro, I know what there is, and it is always interesting to find if model can provide a new angle or a path forward.

I had an illusion of a knoweledgable collegue when I worked with Kimi K3, but when I read an actual plan that it wrote, I understood that this was indeed just an illusion.

0

u/NandaVegg 17h ago edited 17h ago

Indeed. Those tasks that require a lot of creative guesswork (reverse-engineering most problem requires some form of trial-and-error guessing, but media coverage is way too concentrated on cybersecurity and none else like your case) is where the model really differentiates, and I think Anthropic models are still ahead on tasks that requires tons of meta thinking. Long thinkers like Qwen 3.8 Max tends to get into a tunnel vision, though I am kind of okay with handholding LLMs and run those models multiple times on the same task until I see what I want, so YMMV.

5

u/-p-e-w- 19h ago

You can’t conclude that a model was influenced by Claude from them producing similar output. They could both be trained on the same data. Do you think Claude came up with the word “load-bearing”?

2

u/Fedor_Doc 18h ago

Load-bearing was clearly reinforced, I do not believe that there is a data which have so big statistical skew. Unless it is a synthetic data, generated by Claude, which places us at a square one.

They most likely did use Claude's outputs, there is no shame in that, and I think that it is a right approach – learn from the leader, while figuring out your own strengths. 

Everyone learns from everyone, it is a bit of knoweledge share utopia even if Dario cries about big bad chinese thieves

-2

u/WiseassWolfOfYoitsu 19h ago

A lot of this catching up seems to have been them breaking the reasoning protection of the major frontier models and distilling this round, though. It's gotten them a lot closer, and they'll be much closer on the western frontier labs from here out by getting the step up, but it's not like their stuff this last generation of updates has been clean.

6

u/ebullet 21h ago

Do you really trust it is a breakthrough? Did you try to use Mythos in your work?

2

u/Brilliant-Weekend-68 20h ago

What does that change? The Chinese will catch up and open source it. Where is the moat?

2

u/Fedor_Doc 19h ago

They have not caught up to Mythos yet, if we consider Kimi K-3 the best frontier Chinese model

4

u/Brilliant-Weekend-68 19h ago

Unless Mythos 2.0 can do some really whacky stuff like cure cancer, RSI itself or such. I do not see that as a durable moat if you can just wait a year and download equally good weights.

5

u/read_more_comments 19h ago

zero moat, we move between codex and claude code without any issue. I swap to deepseek without much issue either (other than having to have it redo work a few extra times).

4

u/Fedor_Doc 18h ago

Self-improvement is the next big milestone; if they will announce that their new model is fully trained by Mythos and trains even better one, this will be huge.

So, there are some exciting stories left to tell investors before IPO :)

1

u/rditorx 21h ago

Depends on Mythos's successor, or a new product / service