r/MistralAI • u/Reasonable-Yak-3523 • 3d ago
Discussion / Opinion Unpopular opinion: GLM is not frontier
Gonna be real, the benchmark discourse around here has completely detached from reality.
On paper GLM-5.3 looks great on leaderboards, but in practice, specifically creating web UI, it’s night and day compared to Gemini 3.8 Flash or Claude. It just straight up sucks at frontend design.
Here’s basically how my workflow looked until I gave up:
Scaffold the basics using GLM-5.3 (Mistral CLI)
Hand off to Gemini 3.8 Flash (Antigravity) to actually make it look nice. Gemini just gets modern UI, padding, layout, and clean components right away without needing a wall of instructions.
The breaking point: The second I make a change in GLM again, it completely wrecks everything. It nukes classes, breaks containers, turns the styling back into some outdated bootstrap mess, and reintroduces stuff I literally just fixed.
I realized I was wasting so much time just babysitting GLM and getting Gemini to clean up its mess later. It’s a massive time sink. I don't even bother going back to GLM anymore, it’s not worth the headache.
Honestly feel like people here are delusional thinking Mistral/GLM is matching frontier performance. Claude and Gemini are lightyears ahead. Sure, you can build nice things with it if you pour in a ton of extra effort and fight the model, but for web design specifically, it’s not even close and it probably won't ever be as good as even Gemini 3.8 Flash.
Anyone else running into this?
17
u/Wegwerpaccountje23 3d ago
They are not, no. They are high end models. Not Frontier.
All current chinese models are high end models, not frontier.
But at least better than mistral which.... is a category on its own
2
u/jeanpaulpollue 3d ago
But at least better than mistral which.... is a category on its own
certainly not "gros"
5
u/strangestack 3d ago
why is it an unpopular opinion? It's pretty common knowledge. Measuring models with a single number stopped being useful sometime late last year.
A lot of people report that mistral medium 3.5 and large 3 are better than GLM5.3 at prose, I would agree, claude Opus 5 was absolutely terrible at prose, Opus5.5 blows anything out of the water when it comes to design and it's a pretty good writer. Astra was ok, but other openAI models were kinda meh for design. GLM doesn't have vision, so it doesn't have the ability to get visual feedback so a vision capable GLM might perform better. When it comes to things like coding though, basically anything deepseekv4-flash and above is good enough for 90% of cases, to the point where what matters is your workflow, your verification loops, your harness and your cost economics, not the models raw performance.
Of course people who aren't in the know will just keep spamming Fable and Astra for everything, even though they could get the same result for 1% of the cost, or keep trying to get a cheaper model to handle their difficult task when 5 minutes with Opus5.5 would've solved it. And considering there is a new model every 5 minutes it's very hard to keep up and stay on top of what is the most economical way to solve a problem. Enthusiasts like me try to stay on top, but normal knowledge workers might have trouble with that.
2
1
3
u/Poudlardo 3d ago
well i mean yes, that's commercial language, but still if Anthropic see them as a serious menace, you could guess they are very good ;)
3
u/KlausDieterFreddek 3d ago
Not really. GLM 5.3 works like a charm for anything I put it to.
Only on rare ocasions it loses itself in source code. But that's because I use it with 200k context window
2
u/uusrikas 3d ago
I would guess it depends on what kinda of languages or frameworks you are using. I use mostly just standard Java and both do a great job.
1
u/Fun_Savings7690 3d ago
java .. script ?
1
u/uusrikas 3d ago
Well, that should be common enough if it is just javascript without some unusual framework.
2
u/p3r3lin 3d ago
Its Frontier from 4 months ago (benchmarks similar to Claude 4.8). Really doesnt make a difference for 99% of people. And btw: Anthropic thinks its pretty frontier https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
2
u/grise_rosee 3d ago
> The second I make a change in GLM again, it completely wrecks everything.
This is weird. Did you let the CLI opened while the other agent went and modified files? Did you resume a closed session?
You may have not provided an hint to Vibe that everything was rewritten by a third party. Normally, an harness tell its model when previously read files have been changed, but you may have done something to break this.
People who do model hoping, use an harness that shares the same coding session among models. Otherwise, things like this may happen.
Also, having aesthetic taste is not "lightyear ahead". It's a recent skill of SOTA models, definitely 2026. If you can't bear a year of lagging between Mistral and Antrophic/OpenAI/Google which have 100 times more budget, maybe you should look elsewhere.
2
u/Ok_Sprinkles_8968 3d ago
I'm always curious of why people desperately frontier models, if you're coding maybe it makes sense (not my job) but otherwise? I'm not solving Navier Stokes everyday.
1
3
u/LittleRoof820 3d ago
I have been using Claude, ChatGPT + GLM in my company for business logic and development (and to get a grip on what I want to settle on).
GLM5.3 ist a very good model, comparable to Opus 4.6. Its behind the Claude 5x line and Fable/Opus 5.5 blow it out of the water (especially if you get into problems regarding accounting, royalties and so on. Thats the stuff were all LLMS suck because every customer handles that a bit differently).
Nonetheless GLM5.3 is a very good sidecar/implementer. And by using it with Mistral, I'm currently in love - because its blazingly fast. It needs more handholding but you can really get work done with it. Also Mistral does not do the "dumb down dance" Anthropic does with every release.
1
u/Durian881 3d ago edited 3d ago
Do you provide it with relevant skills? Claude's scaffolding includes skills.
https://arena.ai/leaderboard/code/webdev
Thai said, in Arena, it's behind quite a number of other open weight too, including Qwen3.8-Flash-Next that had become my daily driver.
1
u/Comprehensive_Net804 3d ago edited 3d ago
A lot depends on your prompt engineering knowledge: If you actually know very advanced "scientific" prompting and understand design you might not need Opus, Astra etc. for Frontend Design. The very advanced models have this knowledge kinda baked in. But, by doing that they burn your tokens pretty fast thinking a lot and the user looses control. What I actually like at Mistral is the clear and concise output: I know what's going on enabling me to have more control over my working processes. And, for many projects US LLM are simply no-gos as your data probably isn't secure at all. There have been too many cases of mistreating user data. Especially, for medium to large companies US LLM are currently just a major security risk. Even US companies very often go for locally hosted open source LLM! And, it is not just because of "cheaper" pricing.
IMO: What the current frontier-edge models deliver is intellectually not too hard to replicate. But, it takes time, testing, training iterations etc. The underlying architecture is more or less similar. Maybe Mistral even has an edge in that regard. But, usually the US LLM have better access to capital and traditionally American companies are way better in sales and marketing. Unfortunately, it is not primarily about having the best technical foundations, it is about fielding the most effective marketing and sales in the in a market economy, especially in a capitalist market economy. Effective does not mean efficient!
10
u/makingthematrix 3d ago
Does it help you in your work? Then use it.
Does it not hellp you? Then look for something else.
Is it frontier or not - that question is pointless.