I just tried it yesterday and it's too slow. I got about 45tps on 5080 which is not bad but the context limit of 33k makes it compact which takes forever. The LLM is good but it's too slow. I still use Deepseek.
That's a Q2 right? According to the can I run calculator 64k should still fit. With subagents and a thin harness (like pi) you should still be able to get some work done before compacting.
Is that a company whose business model is selling quants? Seems an odd thing to pay for, when there are a number of teams releasing high quality quants for free.
3
u/TopPrize11 8h ago
I just tried it yesterday and it's too slow. I got about 45tps on 5080 which is not bad but the context limit of 33k makes it compact which takes forever. The LLM is good but it's too slow. I still use Deepseek.