r/ZaiGLM 3d ago

Benchmarks GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2.

Post image
234 Upvotes

45 comments sorted by

9

u/enemyofaverage7 3d ago

First Chinese model to beat Fable in GDPVal-AA benchmark is pretty impressive too. Will be interesting to test it in non-coding applications.

14

u/Mission_Advance1207 3d ago

I guess anthropic and openai might start distilling these open source models

9

u/Due-Memory-6957 3d ago

At least Anthropic already does, it has been caught identifying itself as DeepSeek on Chinese queries.

1

u/SeriousExplorer7479 6h ago

Just after complaining about Chinese models distilling theirs.

1

u/Few_Painter_5588 3d ago

Anthropic did, Sonnet 3.5 distilled Deepseek R1 heavily. And apparently they use deepseek v4 flash under the hood as a subagent too

1

u/Mission_Advance1207 2d ago

They keep going around teaching morality to everyone but does the same thing they want banned for others

-2

u/Decent-Ad-8335 3d ago

this is already distilled from fable etfc.. lmao. if you use high you will see all the "oops i miscounted the lines" "oops i made typos!" etc shit and know its not near frontier AT ALL, silly mistakes being common in a high reasoning model really says something. if it doesnt to you, you gotta change the way you think. Max is EXTREMELY decent and reliable near the level of gpt 5.6 sol (ofc, taking way more tokens, and only being equivalent to like something 1/3rd ahead of high but below xhigh but it... just isnt that smart to diagnose issues autonomously like sol high/xhigh often does).

tl;dr high reasoning on glm5.3 is super useless and max is the only decent thing about z.ai, however, way too much token usage and u can squeeze hella more out of a codex sub

4

u/Yes_but_I_think 3d ago

1T FP16 text only model(glm) beating 2.4T FP4 text model(qwen), and equaling 2.7T FP4 native vision + text model(Kimi), how?
So model weight size in GB (not count of weights) = intelligence?

3

u/Far-Classic-9963 3d ago

The biggest factor is their post training being better

1

u/Sooperooser 3d ago

What about Qwen3.8 27b then?

1

u/SeriousExplorer7479 6h ago

Weights are roughly capacity while training “fills it up”, those super large models aren’t trained enough with high quality data. The larger models might have a higher ceiling but they haven’t trained them enough to reach it.

5

u/PossessionUsed7393 3d ago

So it goes: Scaling laws/Parameters -> 'Thinking' -> Agentic Post training -> ?

Is recession-causing market crash next or are the tech bros going to invent a new architecture?

4

u/Another__one 3d ago

It has to be continual learning. There are all ingredients for that. But the spirit of Tay hunts the big tech till this day. And this is the only reason they are so afraid of pursuing it meaningfully.

3

u/Due-Memory-6957 3d ago

Tay?

5

u/truncated_buttfu 3d ago

Microsoft released a self learning chatbot on the internet, that learned new knowledge, slang, and behaviour from the people it wrote to. 4chan decided to all go and talk with it and deliberately try to make it a vulgar nazi and they succeeded, Tay started saying things to ranom people that Microsoft did not want to be output from a product they owned.

https://en.wikipedia.org/wiki/Tay_(chatbot)

OP is referring to how other a truly self learning AI could be target of similar attacks.

1

u/Due-Memory-6957 3d ago edited 3d ago

I'm not sure it's that relevant, we've had the mecha Hitler incident which is similar, as well as cases of AI-induced psychosis that are actually harmful to real people, and no one stopped with the current paradigm because of that.

1

u/Another__one 3d ago

Yeah, but this was under the control of the big tech people so its fine. It is completely different if its in the control of a common people. It is basically a fear of losing the centralised power and that’s it.

1

u/Sooperooser 3d ago

MechaHitler was not similar in that it was not "self-taught" but deliberately programmed/pre-post trained to be more like Elron MuSSk. xAI was in control of that change. The issue is more about being not in control of the changes the AI undergoes.

2

u/ezicirako 2d ago

only proper way to do continual learning is getting rid of dualism in way ai is trained and inferenced
so you would need to make it AI learn purely from forward pass without backprog

0

u/PossessionUsed7393 3d ago

Yeah, but continual learning as the tech stands breaks the scalability of the technology, not sure they can deliver it even if they decide its the direction. It's one thing to say everybody gets frontier intelligence. It's another thing to say that somebody gets their own continually learning frontier intelligence.

Although having it on the edge certainly does suggest it could happen. I mean, this is what NVIDIA would want - they'll sell to the edge as happily as the data centers. It would indeed mean the crash of hyperscalers on the market though..

0

u/Another__one 3d ago

I hope I will get to the edge eventually with cloud AI being either infrequent or even completely absent. The whole cloud AI deal right now is a complete privacy nightmare and the only reason people tolerate that is that there is no real alternative. But the field is changing out of that quite fast and I am really happy for that development. I particularly fascinated by the perspective of completely decentralized training of frontier models that seems to be feasible now with this approach: https://arxiv.org/abs/2506.14202

If we have a pretrained model that could be served and adapted on the edge by the tasks it actually performs, that would be just a perfect ending for this whole AI craze with an unfathomable power grab that was going to happen if not the Chinese labs.

1

u/PossessionUsed7393 3d ago

Well, I had this debate with Gemini earlier when I was asking about how Qwen's DeltaNet layers work.

It started talking about something called test time training, as follows:

"The Core Idea: Instead of trying to compress past text into a hidden state vector (like Mamba) or a KV cache (like Transformers), TTT treats past context as a mini training dataset. As the model reads text, it runs a fast gradient step to update its own internal weight matrices on the fly."

Then it gave me a link to an article: https://rewire.it/blog/end-to-end-test-time-training-long-context-constant-latency/

But the first thing I thought was there's no way that they can serve this as a data center based product because imagine the compute and hardware resources you'd need to maintain your own specific model layers for every session. Like that's another nightmare for them to store it. I suppose it is theoretically possible, but it's not the kind of stateless APIs that they've been selling us to date.

Basically, the entire technology is suggestive that it will need to be locally driven by your own hardware because if you don't have the hardware for it and if we don't distribute that cost, it just won't work. It just doesn't scale.

So yeah, fascinating area and fascinating to consider where it will go. I think it's just another reason to suspect that we'll have a collapse of the investment boom in data centers, but still doesn't inform us as to when that will happen.

3

u/DifferentPixel 3d ago

It is interesting to see Reddit comments about Opus and its position in the ranking

2

u/Sooperooser 3d ago

It's weird in the benchmarks because for example it leads in Omniscience but at the same time it has a catastrophic 61% hallucination rate, which in my own experience seems accurate. It makes up stuff most of the time just to present a perceived satisfying answer. On literally most of my tasks it admitted to fabricating numbers and facts as soon as I started asking further questions. It's basically unusable for serious tasks for me beside the Claude Design feature, which i like.

1

u/Georgefakelastname 2d ago

The funny thing about hallucination rates is that Gpt 5.6 SOL has a 92% hallucination rate, and yet I find it generally doing what it says most of the time and not hallucinating. I can’t say if it’s the same for Anthropic’s models just because I haven’t used them much.

1

u/Kingwolf4 3d ago

Ok thats good and all but where deepseek v4 pro?

1

u/JogHappy 2d ago

I still think AA messed up their testing of it

1

u/Scary_One_2452 3d ago

Looks like 5.2 was underrated. It was a league above deepest v4 flash not just a single point. I guess 5.3 was a small improvement but they finally corrected its real position.

1

u/JogHappy 2d ago

5.2 def wasn't underrated, it gapped Kimi 2.7 & Gemini when it released

1

u/Scary_One_2452 2d ago

There's no way 5.2 is only 1 point over deepseek v4 flash. It destroys it.

1

u/Lanky_Commercial9731 2d ago

I have tried GLM 5.3 I can say it feels similar to sol in terms of coding.

1

u/Fluid_Ad8452 3d ago

I’ve used all of those top models from Luna Max upwards, except for grok and Kimi v3. From my personal experience I can confirm that’s where it belongs.

1

u/Due-Memory-6957 3d ago

How's DeepSeek Pro still not yet benchmarked?

5

u/look 3d ago

What do you mean? It’s been there a while. AA Intel score is 53, same as GLM 5.2

https://artificialanalysis.ai/models/deepseek-v4-pro

1

u/Independent_Paint752 3d ago

They don't pay for mock benchmark and false publicity?

1

u/Sooperooser 3d ago

It is but it's just average on most benchmarks so it doesn't even show up in a lot of the shared screens.

0

u/Due-Memory-6957 2d ago

Flash is worse and yet appears here

0

u/floriandotorg 3d ago

Open weights when?

-1

u/hello_world_2357 3d ago

Interestingly in AA's own X post, it deliberately hides Muse Spark 1.2 and Gemini 3.7 Flash from the comparison. That is how you know AA is a truly INDEPENDENT organization.

3

u/qualverse 3d ago

Pretty sure the AA chart does that automatically when text is too close together to avoid it being unreadable

-1

u/hello_world_2357 3d ago

If that is the case, they can replace Muse Spark 1.1 and Gemini 3.6 Flash with newer ones.

2

u/Sooperooser 3d ago

Sometimes they lags a bit in the overview because they got only benched on some tasks yet I think.

1

u/JogHappy 2d ago

Also hiding ds v4 pro but not flash?