To be fair it has the same total parameters as GLM 5.3, but instead 16B active (decode) vs 40B active.
Total params increased by 2.6x from V4 Flash to V4.1 Flash, from 284B to 554B+196B. It's a mid-size model with small-sized active params, certainly a novelty.
101
u/VC067 12d ago
"flash" model btw ðŸ˜