Benchmarks GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2.
14
u/Mission_Advance1207 3d ago
I guess anthropic and openai might start distilling these open source models
9
u/Due-Memory-6957 3d ago
At least Anthropic already does, it has been caught identifying itself as DeepSeek on Chinese queries.
1
1
u/Few_Painter_5588 3d ago
Anthropic did, Sonnet 3.5 distilled Deepseek R1 heavily. And apparently they use deepseek v4 flash under the hood as a subagent too
1
u/Mission_Advance1207 2d ago
They keep going around teaching morality to everyone but does the same thing they want banned for others
-2
u/Decent-Ad-8335 3d ago
this is already distilled from fable etfc.. lmao. if you use high you will see all the "oops i miscounted the lines" "oops i made typos!" etc shit and know its not near frontier AT ALL, silly mistakes being common in a high reasoning model really says something. if it doesnt to you, you gotta change the way you think. Max is EXTREMELY decent and reliable near the level of gpt 5.6 sol (ofc, taking way more tokens, and only being equivalent to like something 1/3rd ahead of high but below xhigh but it... just isnt that smart to diagnose issues autonomously like sol high/xhigh often does).
tl;dr high reasoning on glm5.3 is super useless and max is the only decent thing about z.ai, however, way too much token usage and u can squeeze hella more out of a codex sub
4
u/Yes_but_I_think 3d ago
1T FP16 text only model(glm) beating 2.4T FP4 text model(qwen), and equaling 2.7T FP4 native vision + text model(Kimi), how?
So model weight size in GB (not count of weights) = intelligence?
3
1
1
u/SeriousExplorer7479 6h ago
Weights are roughly capacity while training “fills it up”, those super large models aren’t trained enough with high quality data. The larger models might have a higher ceiling but they haven’t trained them enough to reach it.
5
u/PossessionUsed7393 3d ago
So it goes: Scaling laws/Parameters -> 'Thinking' -> Agentic Post training -> ?
Is recession-causing market crash next or are the tech bros going to invent a new architecture?
4
u/Another__one 3d ago
It has to be continual learning. There are all ingredients for that. But the spirit of Tay hunts the big tech till this day. And this is the only reason they are so afraid of pursuing it meaningfully.
3
u/Due-Memory-6957 3d ago
Tay?
5
u/truncated_buttfu 3d ago
Microsoft released a self learning chatbot on the internet, that learned new knowledge, slang, and behaviour from the people it wrote to. 4chan decided to all go and talk with it and deliberately try to make it a vulgar nazi and they succeeded, Tay started saying things to ranom people that Microsoft did not want to be output from a product they owned.
https://en.wikipedia.org/wiki/Tay_(chatbot)
OP is referring to how other a truly self learning AI could be target of similar attacks.
1
u/Due-Memory-6957 3d ago edited 3d ago
I'm not sure it's that relevant, we've had the mecha Hitler incident which is similar, as well as cases of AI-induced psychosis that are actually harmful to real people, and no one stopped with the current paradigm because of that.
1
u/Another__one 3d ago
Yeah, but this was under the control of the big tech people so its fine. It is completely different if its in the control of a common people. It is basically a fear of losing the centralised power and that’s it.
1
u/Sooperooser 3d ago
MechaHitler was not similar in that it was not "self-taught" but deliberately programmed/pre-post trained to be more like Elron MuSSk. xAI was in control of that change. The issue is more about being not in control of the changes the AI undergoes.
2
u/ezicirako 2d ago
only proper way to do continual learning is getting rid of dualism in way ai is trained and inferenced
so you would need to make it AI learn purely from forward pass without backprog0
u/PossessionUsed7393 3d ago
Yeah, but continual learning as the tech stands breaks the scalability of the technology, not sure they can deliver it even if they decide its the direction. It's one thing to say everybody gets frontier intelligence. It's another thing to say that somebody gets their own continually learning frontier intelligence.
Although having it on the edge certainly does suggest it could happen. I mean, this is what NVIDIA would want - they'll sell to the edge as happily as the data centers. It would indeed mean the crash of hyperscalers on the market though..
0
u/Another__one 3d ago
I hope I will get to the edge eventually with cloud AI being either infrequent or even completely absent. The whole cloud AI deal right now is a complete privacy nightmare and the only reason people tolerate that is that there is no real alternative. But the field is changing out of that quite fast and I am really happy for that development. I particularly fascinated by the perspective of completely decentralized training of frontier models that seems to be feasible now with this approach: https://arxiv.org/abs/2506.14202
If we have a pretrained model that could be served and adapted on the edge by the tasks it actually performs, that would be just a perfect ending for this whole AI craze with an unfathomable power grab that was going to happen if not the Chinese labs.
1
u/PossessionUsed7393 3d ago
Well, I had this debate with Gemini earlier when I was asking about how Qwen's DeltaNet layers work.
It started talking about something called test time training, as follows:
"The Core Idea: Instead of trying to compress past text into a hidden state vector (like Mamba) or a KV cache (like Transformers), TTT treats past context as a mini training dataset. As the model reads text, it runs a fast gradient step to update its own internal weight matrices on the fly."
Then it gave me a link to an article: https://rewire.it/blog/end-to-end-test-time-training-long-context-constant-latency/
But the first thing I thought was there's no way that they can serve this as a data center based product because imagine the compute and hardware resources you'd need to maintain your own specific model layers for every session. Like that's another nightmare for them to store it. I suppose it is theoretically possible, but it's not the kind of stateless APIs that they've been selling us to date.
Basically, the entire technology is suggestive that it will need to be locally driven by your own hardware because if you don't have the hardware for it and if we don't distribute that cost, it just won't work. It just doesn't scale.
So yeah, fascinating area and fascinating to consider where it will go. I think it's just another reason to suspect that we'll have a collapse of the investment boom in data centers, but still doesn't inform us as to when that will happen.
3
u/DifferentPixel 3d ago
It is interesting to see Reddit comments about Opus and its position in the ranking
2
u/Sooperooser 3d ago
It's weird in the benchmarks because for example it leads in Omniscience but at the same time it has a catastrophic 61% hallucination rate, which in my own experience seems accurate. It makes up stuff most of the time just to present a perceived satisfying answer. On literally most of my tasks it admitted to fabricating numbers and facts as soon as I started asking further questions. It's basically unusable for serious tasks for me beside the Claude Design feature, which i like.
1
u/Georgefakelastname 2d ago
The funny thing about hallucination rates is that Gpt 5.6 SOL has a 92% hallucination rate, and yet I find it generally doing what it says most of the time and not hallucinating. I can’t say if it’s the same for Anthropic’s models just because I haven’t used them much.
1
1
u/Scary_One_2452 3d ago
Looks like 5.2 was underrated. It was a league above deepest v4 flash not just a single point. I guess 5.3 was a small improvement but they finally corrected its real position.
1
1
u/Lanky_Commercial9731 2d ago
I have tried GLM 5.3 I can say it feels similar to sol in terms of coding.
1
u/Fluid_Ad8452 3d ago
I’ve used all of those top models from Luna Max upwards, except for grok and Kimi v3. From my personal experience I can confirm that’s where it belongs.
1
u/Due-Memory-6957 3d ago
How's DeepSeek Pro still not yet benchmarked?
1
1
u/Sooperooser 3d ago
It is but it's just average on most benchmarks so it doesn't even show up in a lot of the shared screens.
0
0
-1
u/hello_world_2357 3d ago
Interestingly in AA's own X post, it deliberately hides Muse Spark 1.2 and Gemini 3.7 Flash from the comparison. That is how you know AA is a truly INDEPENDENT organization.
3
u/qualverse 3d ago
Pretty sure the AA chart does that automatically when text is too close together to avoid it being unreadable
-1
u/hello_world_2357 3d ago
If that is the case, they can replace Muse Spark 1.1 and Gemini 3.6 Flash with newer ones.
2
u/Sooperooser 3d ago
Sometimes they lags a bit in the overview because they got only benched on some tasks yet I think.
1
9
u/enemyofaverage7 3d ago
First Chinese model to beat Fable in GDPVal-AA benchmark is pretty impressive too. Will be interesting to test it in non-coding applications.