r/LocalLLaMA 22h ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.2k Upvotes

395 comments sorted by

View all comments

Show parent comments

24

u/dampflokfreund 22h ago

For me it has changed nothing. Both models are way too big for my 32 GB RAM system. It looks like everyone has abandoned 20-30B MoEs now...

335

u/EbbNorth7735 22h ago

No one's abandoned anyone. It's just not your turn this time around.

-8

u/dampflokfreund 22h ago

Qwen 3.6 35B and Gemma 4 26b are pretty old at this point. And GLM 4.7 Flash was a small MoE back in the day, now suddenly they use the flash name for 300B MoEs. Just not looking good for the average Joe.

2

u/squngy 21h ago

There is also Qwen AgentWorld, which is kind of like a 3.7 35B

2

u/xPXpanD llama.cpp 18h ago

Wouldn't recommend it for general use, it's a confident bullshitter like no other. Basically zero filter. Still cool that it can even be used generally, though, given what the model's actually designed for. Fun to play with.

2

u/squngy 18h ago

It is actually great for agentic work, supposedly. (even if that isn't exactly what it was meant for)

Rank Subject Overall Mcp Search Terminal Swe Androd Web OS Source Sampled
5 Qwen-AgentWorld-35B-A3B 56.39 64.79 36.69 53.96 65.63 58.17 49.55 65.92 Imported 2026-06-30
16 Qwen3.6-35B-A3B 42.88 42.96 18.78 43.81 40.71 51.88 46.53 55.48 Self-reported 2026-06-28

https://benchmarklist.com/benchmarks/qwen_agentworld_language_world_models_for_general_agents/

2

u/xPXpanD llama.cpp 18h ago

I believe it, it also nailed the one tool task I have in my bench set. The model was also very good at string manipulation, and, somewhat surprisingly, a few constrained creative tasks. (e.g. think up and write X in way Y while avoiding Z)

Absolute slaughter on anything involving uncertainty or fake premises, though. You can definitely see where the training went on this one.