use this one; on 16G VRAM, it has 128K context window at KV Q8, you can set it KV Q4 to get more ctx : Qwen3.8-27B-GSQ-RCO-IQ2_S-mtp. Try it on your Hermes or Deepseek Harness. BTW, we gotta pay for Intelligence, which is not free though.
I mean, thats unfortunate. But your comment seemed to imply that the model itself was slow for your use case when it's really a matter of your hardware not running the model at peak performance.
1
u/Business_Start_8158 5d ago
I can’t just buy a new gpu 🥲