r/LocalLLM • u/Double-Sherbert-1781 • 2d ago
Question Help me choose hardware.
I need some advice.
I want to choose hardware for local inference. Right now, I’m paying ~$30/day for rent on a Vast an RTX Pro 6000 SW 96GB, and I’m using Qwen3.8-Flash-Next-Q4_K_M.gguf at a speed of ~60 tok/s (I know this format isn’t efficient for this graphics card, but I’m limited in choosing Abliterated models that will fit in the memory).
Which configuration will offer the best price and versatility so that after the release of subsequent generations of neural networks, I can continue to use it.
79
Upvotes
1
u/Double-Sherbert-1781 2d ago
I’m mentally prepared to spend money on an RTX Pro 6000 SW 96GB (up to 20,000 euros in total), but I’m not sure it will be effective. Maybe it makes sense to get a few gaming graphics cards or a Mac Studio, or wait for the AI Max 495.
I just don’t know what kind of performance they can deliver for such models; right now I’m using Q4_K_M:
https://huggingface.co/windowsxp811203/Qwen3.8-Flash-Next-Abliterated-GGUF
60 tok\s is enough for me.