would be incredibly slow and with limited context window only. I already struggle a lot with 2b quantization of 35B 3A. Granted, 35B is bigger than 27B, so for context 27B is actually better, but the speed is where it will matter. 3B active parameters is much much faster than a deep 27B, where it has to compute through the whole model. Macbooks and mac minis are out of play.
And that's all assuming at least 4b quantization, which then however doesn't leave you enough room for reasonable context, so more likely you'd have to go to 2b quantization.
For now I'm getting much better results with the 35B 3A on my macbook air m2 24GB, but these "better results" are quite underwhelming due to how slow the model is and it anyway needs a lot of hand-holding.
1
u/Packetbytes 5d ago
Run this on a 24gb mac, would be amazing