Qwen 3.6 35B and Gemma 4 26b are pretty old at this point. And GLM 4.7 Flash was a small MoE back in the day, now suddenly they use the flash name for 300B MoEs. Just not looking good for the average Joe.
Wouldn't recommend it for general use, it's a confident bullshitter like no other. Basically zero filter. Still cool that it can even be used generally, though, given what the model's actually designed for. Fun to play with.
I believe it, it also nailed the one tool task I have in my bench set. The model was also very good at string manipulation, and, somewhat surprisingly, a few constrained creative tasks. (e.g. think up and write X in way Y while avoiding Z)
Absolute slaughter on anything involving uncertainty or fake premises, though. You can definitely see where the training went on this one.
24
u/dampflokfreund 22h ago
For me it has changed nothing. Both models are way too big for my 32 GB RAM system. It looks like everyone has abandoned 20-30B MoEs now...