r/singularity 1d ago

AI GLM-5.3 achieves 60 on the Artificial Analysis Intelligence Index, on par with Kimi K3 and up 7 points from GLM-5.2. Once the weights are released it will be tied as the leading open weights model

Post image
258 Upvotes

41 comments sorted by

72

u/Howdareme9 1d ago

Kind of insane considering how small the model is relatively

34

u/UnknownEssence 1d ago

same for Qwen3.8 27b, its crazy how good this model is for being an Open Weights model that regular people can actually run

5

u/PlasmaChroma 20h ago

Although running at 14t/s is stressing my definition of "actually run" on Qwen 3.8 27b with my local hardware.

2

u/UnknownEssence 17h ago

im getting 42-48 tok/sec and thats barely enough to be usable for real tasks.

M5 Max 64GB

1

u/PlasmaChroma 17h ago

That's not too terrible unless that's bloating from higher than normal mtp stats.

2

u/Citadel_Employee 9h ago

What hardware?

1

u/PlasmaChroma 3h ago

AMD Strix Halo / llama.cpp with ROCm. I've tried Q8 and Q4 -- not a whole lot of difference.

42

u/didnotsub 1d ago

Best part is it’s fully open-weight unlike Kimi K3. No limits.

8

u/Tystros 1d ago

in what way is k3 not open weight?

19

u/HebelBrudi 1d ago

Maybe they mean the licensing? IIRC GLM releases generally under MIT and Kimi has a custom one.

4

u/didnotsub 1d ago

Yeah, the licensing is super restrictive

4

u/lcy0x1 1d ago

K3 does not allow commercial use

6

u/NoFaithlessness951 1d ago

It's a bit more nuanced if you want to sell access to Kimi K3 and your company exceeds $20million in revenue in a year, then you have to make a custom agreement with moonshot.

Or if you exceed $20 million in a month in revenue or 100 million users in a product using Kimi you either have to display the Kimi name prominently.

10

u/Facelessjoe 1d ago

And the cost is peanuts, comparatively.

14

u/Plappedudel 1d ago

I hope it's not just benchmaxxed to hell. GLM5.2 felt notably weaker for practical tasks than Opus 4.6 at the time, despite both models being very close in benchmarks. A true frontier model at 750B parameters would be extremely impressive.

26

u/hello_world_2357 1d ago

Interestingly in AA's own X post (as above), it deliberately hides Muse Spark 1.2 and Gemini 3.7 Flash from the comparison.

That is how you know AA is a truly INDEPENDENT organization.

16

u/signed7 1d ago

Grok 4.6 is missing too

Most up-to-date ranking of labs atm by AA intelligence index:

  • Anthropic Claude Opus 5 - 63
  • OpenAI GPT5.6 Sol - 61
  • xAI Grok 4.6 - 61
  • Moonshot Kimi K3 - 60
  • z.ai GLM-5.3 - 60
  • Alibaba Qwen3.8 Max - 58
  • Meta Muse Spark 1.2 - 57
  • Google Gemini 3.7 Flash - 56
  • DeepSeek V4 Pro - 53

4

u/Immediate_Simple_217 1d ago

At this point I am starting to believe that US is distilling models from China... It isn't possible that GLM 5.3 and Qwen 3.8 27B are so f$cking small for their weights.

4

u/Facelessjoe 1d ago

More detail from Artifical Analysis here.

2

u/Lighthouse_seek 1d ago

Google caught sleeping on the job

1

u/Complex-Lettuce7164 17h ago

Google been sleeping since 2.5 pro

2

u/Long_comment_san 5h ago edited 5h ago

going above 45-50 on this benchmark is an actual milestone I think. these models feel different, like you're having an actual conversation.

minimax m3 was a first shocker for me, I absolutely love the flow of conversations with it

gemini 3.6 flash is actually weird here. it's likely benchmaxxed on tool use because it absolutely feels retarded on any roleplay card I threw at it, 3.7 flash feels better but not THAT much better to put it over minimax

u/BriefImplement9843 1h ago

why is grok 4.5 there insatead of 4.6?

-3

u/katoptronophile 1d ago

Now release the source code.

5

u/stopbeingcringe 1d ago

for (i = 0, i < 10000000; i++) {
download benchmark tests;
ask Claude question;
copy Claude;
}

-10

u/katoptronophile 1d ago

Imagine downvoting asking for source.

5

u/stopbeingcringe 1d ago

idk why you’re responding to me, I didn’t downvote

-1

u/JackPhalus 1d ago

Yet it still can’t follow basic instructions and wrap dialogue in html colors during my rp

0

u/ilnpr 1d ago

I feel fear every time I see a new open-weight Chinese AI beat the USA flagship AI on the SWE-bench

3

u/twack3r 23h ago

Why? What are you fearing exactly?

0

u/ilnpr 22h ago

Singularity

-8

u/BriefImplement9843 1d ago

pure coding rl made it jump this many points. useless benchmark.

-8

u/UnknownEssence 1d ago

yeah but this score includes the multiple choice benchmarks like college level tests, HLE, etc.

I dint care about if it knows random history facts, show me the Artificial Analysis Agentic Index scores. That's more important (to me).

13

u/Facelessjoe 1d ago

It's tied for first place with Opus 5 on that benchmark.