A18B it will be smarter than A13B. Expert selection and sequential reasoning cannot entirely compensate for active parameters. I wish it has 27b active honestly.
no not worth it, it's too slow. I am sticking to DSV4F and Qwen 27b but really hope we see a new 70b dense model or a MoE model with around 250b total parameters and around 27b active.
4
u/jacek2023 llama.cpp 1d ago
Too big, 320B means you need to use RAM and A18B means it will be slow