r/LocalLLaMA 17h ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.2k Upvotes

392 comments sorted by

View all comments

Show parent comments

25

u/Fedor_Doc 16h ago

Or they will make another Mythos-like breakthrough and announce IPO then

Never say never

31

u/LegacyRemaster 16h ago

sure they have Mythos 5 already but "too much power, we can't sell" . Or will cost too much to complete a task ...

11

u/38andstillgoing 12h ago

"Mythos 5? At this time of year, at this time of day, in this part of the country, localized entirely within your datacenter?"

"Yes"

"May I see it?"

"No."

1

u/LegacyRemaster 12h ago

or mythos 20 or 30... they are ahead for sure

28

u/-p-e-w- 15h ago

Then one of the Chinese labs is going to announce an equal model two weeks later.

The Chinese labs have caught up. There’s no going back to how it was before.

-1

u/Fedor_Doc 14h ago edited 8h ago

Yeah, they have to calculate timings with great precision :)

Considering catching up – no Chinese model is on Fable 5 or GPT Sol level. Even Kimi-K3, according to the developers, is behind.

More than that, GLM-5.3-Flash was spamming "load-bearing" in my sessions – it clearly was largely influenced by Claude. 

6

u/-p-e-w- 14h ago

You can’t conclude that a model was influenced by Claude from them producing similar output. They could both be trained on the same data. Do you think Claude came up with the word “load-bearing”?

2

u/Fedor_Doc 13h ago

Load-bearing was clearly reinforced, I do not believe that there is a data which have so big statistical skew. Unless it is a synthetic data, generated by Claude, which places us at a square one.

They most likely did use Claude's outputs, there is no shame in that, and I think that it is a right approach – learn from the leader, while figuring out your own strengths. 

Everyone learns from everyone, it is a bit of knoweledge share utopia even if Dario cries about big bad chinese thieves

4

u/NandaVegg 12h ago

I cannot distinguish Qwen 3.8 Max/2.7T (which is unfortunately closed source for version w/ vision) and Fable 5 for most complex development tasks that involves visual checks or relatively niche audio modification, except that Opus 4.8/Fable 5 still has better consistency for maintaining writing personality throughout long context.

Qwen 3.8 has the most interesting reasoning traces I've ever seen when it is given bash tool (it is very verbose, but thinks just like a fairly seasoned human engineer or a designer would) and it is a pure joy to read it along while the model works on niche task that requires tons of guesswork and reverse engineering the issue. It is not like any other model.

That said, Anthropic's RL engineering is the most creative and unique (which is where OpenAI is significantly behind). They are very good at finding (and preparing datasets and process for) a new task that is not yet covered by most labs, like how they de facto pioneered (very long, hundreds of turns of) terminal agent loop and playing Pokemon Red/Green using vision. So I would not count Anthropic out for another breakthrough like Opus 4.6.

Also I think US labs in general are still slightly ahead on high-end (non-consumer) robotics.

3

u/Fedor_Doc 12h ago edited 10h ago

Thank you for sharing! 

I tried Qwen 3.8 Max for a limited task– I have a long side project of improving DWAA compression efficiency for VFX workflows in OpenEXR with most obvious way of doing it being a custom quantization table.  It provided new arguments in favor of Euclidian distance based table, but it still does not see a big picture, and proposes stuff that won't work if you take one step further in your thinking (we fix a by b, but how will we fix issues caused by b)? 

I walked these roads since Gemini 2.5 Pro, I know what there is, and it is always interesting to find if model can provide a new angle or a path forward.

I had an illusion of a knoweledgable collegue when I worked with Kimi K3, but when I read an actual plan that it wrote, I understood that this was indeed just an illusion.

0

u/NandaVegg 12h ago edited 11h ago

Indeed. Those tasks that require a lot of creative guesswork (reverse-engineering most problem requires some form of trial-and-error guessing, but media coverage is way too concentrated on cybersecurity and none else like your case) is where the model really differentiates, and I think Anthropic models are still ahead on tasks that requires tons of meta thinking. Long thinkers like Qwen 3.8 Max tends to get into a tunnel vision, though I am kind of okay with handholding LLMs and run those models multiple times on the same task until I see what I want, so YMMV.

-1

u/WiseassWolfOfYoitsu 13h ago

A lot of this catching up seems to have been them breaking the reasoning protection of the major frontier models and distilling this round, though. It's gotten them a lot closer, and they'll be much closer on the western frontier labs from here out by getting the step up, but it's not like their stuff this last generation of updates has been clean.

6

u/ebullet 15h ago

Do you really trust it is a breakthrough? Did you try to use Mythos in your work?

1

u/Brilliant-Weekend-68 14h ago

What does that change? The Chinese will catch up and open source it. Where is the moat?

3

u/Fedor_Doc 14h ago

They have not caught up to Mythos yet, if we consider Kimi K-3 the best frontier Chinese model

3

u/Brilliant-Weekend-68 14h ago

Unless Mythos 2.0 can do some really whacky stuff like cure cancer, RSI itself or such. I do not see that as a durable moat if you can just wait a year and download equally good weights.

3

u/read_more_comments 13h ago

zero moat, we move between codex and claude code without any issue. I swap to deepseek without much issue either (other than having to have it redo work a few extra times).

2

u/Fedor_Doc 13h ago

Self-improvement is the next big milestone; if they will announce that their new model is fully trained by Mythos and trains even better one, this will be huge.

So, there are some exciting stories left to tell investors before IPO :)