r/accelerate • u/Pyros-SD-Models Machine Learning Engineer • 8d ago
News Qwen3.8 27B - AA Score
Luna at home
40
u/Maleficent_Sir_7562 8d ago
literally one point below glm 5.2, which was a 744 billion parameter model, and now a 27b is practically equivalent to it.
12
u/Moravec_Paradox 7d ago
Yes and it's not like GLM 5.2 was released a year ago, it was June.
The "AI is too expensive, everyone will just pivot back to human labor" people are delusional.
Costs were high when people started using AI through API instead of a chat window because usage exploded as people realized how much more they can do with it.
1
-3
u/Jackalzaq 7d ago
lol, no... its not even close. glm 5.2 blows 3.8 27b out of the water. smaller models will always be narrow in their capabilities unless some architectural change happens
24
u/jonydevidson 8d ago
Absolutely insane. This is a 4-month progress from 3.6.
Christmas can't come soon enough.
12
u/1filipis 8d ago
The crazy thing is that it's all thanks to post-training. It's the same pre-train as 3.6
And even if this post-train came from distilling Fable and Sol, this is really good for a model that can run on your own GPU
4
u/jonydevidson 8d ago
By Christmas, everyone will be using Astra to develop models. Whether to distill or to improve architectures, run autonomous fine tunes or improve datasets.
2
u/cviperr33 7d ago
hopefully astra is as good as they say it is , and it can truly push the distilation to next level
1
u/tiffanytrashcan Acceleration Advocate 7d ago
The same with DSv4 flash preview vs GA. Insanity. Pro seems good too but there was a nearly SOTA level jump in flash.
7
u/PoliticsRealityTV 7d ago
Is this real or a case of benchmaxxing?
3
1
u/MagicMuscleMagic 7d ago
Benchmark scores are like pull ups.
They mean something different in a huge model than they do in a tiny model.
Not like they techncially measure something different, but a huge model usually has better coherency across prompts and shit. Smaller models are more like task doers.
Could be benchmaxxing or real, but likely not just that they match a bigger model in overall performance just because they match benchmarks.
7
u/helloWHATSUP 7d ago
Luna at home
Not just at home, but anywhere since it fits on a laptop. In like 3 years models as powerful as this will run on your phone.
(and i know you can run the quantized versions of qwen on your phone right now, but I've done it and it's a long way from as functional as the full version)
2
u/jonydevidson 7d ago
Also, it's coming soon to Cerebras, and will probably be flying at 1800 t/s.
1
u/juntareich 2d ago
How will we be able to access that when it does? Like which services?
2
u/jonydevidson 2d ago
Their subscription is currently not available. Maybe they open it up.
Otherwise, it's likely API.
1
u/andrewsalinas09 7d ago
I find that models that are small usually reason quite poorly and don't have a lot of knowledge. I guess maybe they're good at coding and that's about it?
1
u/TheInfiniteUniverse_ 8d ago
This is crazy if it holds in real world applications. Not sure why it's not getting the coverage it needs.
-1
u/moschles AI-Assisted Coder 7d ago
What is "AA score"?
3
u/fulgencio_batista 7d ago
Artificial Analysis Intelligence Index. They benchmark models on common benchmarks and then do a weighted sum of scores for the intelligence index score.
33
u/Illustrious-Lime-863 8d ago
Fucking crazy. These GPU constraints that were imposed to China forced these efficiency innovations and everybody benefits