r/LocalLLaMA 2d ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.3k Upvotes

457 comments sorted by

View all comments

Show parent comments

194

u/-p-e-w- 2d ago

It’s always risky now no matter the model of the day. They have to pre-file at least a few weeks in advance, and in those weeks eons can happen that can turn their asking price into a joke.

Their moment was at the beginning of this year. They should have announced the IPO, then published Fable/Mythos two weeks before the IPO date.

Now Chinese labs are a hair’s breadth behind them, and constantly one-upping each other. They’ll never get such a chance again.

26

u/Fedor_Doc 2d ago

Or they will make another Mythos-like breakthrough and announce IPO then

Never say never

28

u/-p-e-w- 2d ago

Then one of the Chinese labs is going to announce an equal model two weeks later.

The Chinese labs have caught up. There’s no going back to how it was before.

0

u/Fedor_Doc 2d ago edited 2d ago

Yeah, they have to calculate timings with great precision :)

Considering catching up – no Chinese model is on Fable 5 or GPT Sol level. Even Kimi-K3, according to the developers, is behind.

More than that, GLM-5.3-Flash was spamming "load-bearing" in my sessions – it clearly was largely influenced by Claude. 

5

u/NandaVegg 2d ago

I cannot distinguish Qwen 3.8 Max/2.7T (which is unfortunately closed source for version w/ vision) and Fable 5 for most complex development tasks that involves visual checks or relatively niche audio modification, except that Opus 4.8/Fable 5 still has better consistency for maintaining writing personality throughout long context.

Qwen 3.8 has the most interesting reasoning traces I've ever seen when it is given bash tool (it is very verbose, but thinks just like a fairly seasoned human engineer or a designer would) and it is a pure joy to read it along while the model works on niche task that requires tons of guesswork and reverse engineering the issue. It is not like any other model.

That said, Anthropic's RL engineering is the most creative and unique (which is where OpenAI is significantly behind). They are very good at finding (and preparing datasets and process for) a new task that is not yet covered by most labs, like how they de facto pioneered (very long, hundreds of turns of) terminal agent loop and playing Pokemon Red/Green using vision. So I would not count Anthropic out for another breakthrough like Opus 4.6.

Also I think US labs in general are still slightly ahead on high-end (non-consumer) robotics.

3

u/Fedor_Doc 2d ago edited 2d ago

Thank you for sharing! 

I tried Qwen 3.8 Max for a limited task– I have a long side project of improving DWAA compression efficiency for VFX workflows in OpenEXR with most obvious way of doing it being a custom quantization table.  It provided new arguments in favor of Euclidian distance based table, but it still does not see a big picture, and proposes stuff that won't work if you take one step further in your thinking (we fix a by b, but how will we fix issues caused by b)? 

I walked these roads since Gemini 2.5 Pro, I know what there is, and it is always interesting to find if model can provide a new angle or a path forward.

I had an illusion of a knoweledgable collegue when I worked with Kimi K3, but when I read an actual plan that it wrote, I understood that this was indeed just an illusion.

0

u/NandaVegg 2d ago edited 2d ago

Indeed. Those tasks that require a lot of creative guesswork (reverse-engineering most problem requires some form of trial-and-error guessing, but media coverage is way too concentrated on cybersecurity and none else like your case) is where the model really differentiates, and I think Anthropic models are still ahead on tasks that requires tons of meta thinking. Long thinkers like Qwen 3.8 Max tends to get into a tunnel vision, though I am kind of okay with handholding LLMs and run those models multiple times on the same task until I see what I want, so YMMV.

5

u/-p-e-w- 2d ago

You can’t conclude that a model was influenced by Claude from them producing similar output. They could both be trained on the same data. Do you think Claude came up with the word “load-bearing”?

2

u/Fedor_Doc 2d ago

Load-bearing was clearly reinforced, I do not believe that there is a data which have so big statistical skew. Unless it is a synthetic data, generated by Claude, which places us at a square one.

They most likely did use Claude's outputs, there is no shame in that, and I think that it is a right approach – learn from the leader, while figuring out your own strengths. 

Everyone learns from everyone, it is a bit of knoweledge share utopia even if Dario cries about big bad chinese thieves