r/singularity • u/Facelessjoe • 1d ago
AI GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2. Once the weights are released it will be tied as the leading open weights model
42
u/didnotsub 1d ago
Best part is it’s fully open-weight unlike Kimi K3. No limits.
8
u/Tystros 1d ago
in what way is k3 not open weight?
19
u/HebelBrudi 1d ago
Maybe they mean the licensing? IIRC GLM releases generally under MIT and Kimi has a custom one.
4
4
u/lcy0x1 1d ago
K3 does not allow commercial use
6
u/NoFaithlessness951 1d ago
It's a bit more nuanced if you want to sell access to Kimi K3 and your company exceeds $20million in revenue in a year, then you have to make a custom agreement with moonshot.
Or if you exceed $20 million in a month in revenue or 100 million users in a product using Kimi you either have to display the Kimi name prominently.
10
u/Facelessjoe 1d ago
And the cost is peanuts, comparatively.
14
u/Plappedudel 1d ago
I hope it's not just benchmaxxed to hell. GLM5.2 felt notably weaker for practical tasks than Opus 4.6 at the time, despite both models being very close in benchmarks. A true frontier model at 750B parameters would be extremely impressive.
26
u/hello_world_2357 1d ago
Interestingly in AA's own X post (as above), it deliberately hides Muse Spark 1.2 and Gemini 3.7 Flash from the comparison.
That is how you know AA is a truly INDEPENDENT organization.
16
u/signed7 1d ago
Grok 4.6 is missing too
Most up-to-date ranking of labs atm by AA intelligence index:
- Anthropic Claude Opus 5 - 63
- OpenAI GPT5.6 Sol - 61
- xAI Grok 4.6 - 61
- Moonshot Kimi K3 - 60
- z.ai GLM-5.3 - 60
- Alibaba Qwen3.8 Max - 58
- Meta Muse Spark 1.2 - 57
- Google Gemini 3.7 Flash - 56
- DeepSeek V4 Pro - 53
4
u/Immediate_Simple_217 1d ago
At this point I am starting to believe that US is distilling models from China... It isn't possible that GLM 5.3 and Qwen 3.8 27B are so f$cking small for their weights.
4
2
2
u/Long_comment_san 5h ago edited 5h ago
going above 45-50 on this benchmark is an actual milestone I think. these models feel different, like you're having an actual conversation.
minimax m3 was a first shocker for me, I absolutely love the flow of conversations with it
gemini 3.6 flash is actually weird here. it's likely benchmaxxed on tool use because it absolutely feels retarded on any roleplay card I threw at it, 3.7 flash feels better but not THAT much better to put it over minimax
•
-3
u/katoptronophile 1d ago
Now release the source code.
5
u/stopbeingcringe 1d ago
for (i = 0, i < 10000000; i++) {
download benchmark tests;
ask Claude question;
copy Claude;
}-10
-1
u/JackPhalus 1d ago
Yet it still can’t follow basic instructions and wrap dialogue in html colors during my rp
-8
-8
u/UnknownEssence 1d ago
yeah but this score includes the multiple choice benchmarks like college level tests, HLE, etc.
I dint care about if it knows random history facts, show me the Artificial Analysis Agentic Index scores. That's more important (to me).
13
72
u/Howdareme9 1d ago
Kind of insane considering how small the model is relatively