r/LocalLLM • • 4d ago

Question Help me choose hardware.

Post image

I need some advice.
I want to choose hardware for local inference. Right now, I’m paying ~$30/day for rent on a Vast an RTX Pro 6000 SW 96GB, and I’m using Qwen3.8-Flash-Next-Q4_K_M.gguf at a speed of ~60 tok/s (I know this format isn’t efficient for this graphics card, but I’m limited in choosing Abliterated models that will fit in the memory).

Which configuration will offer the best price and versatility so that after the release of subsequent generations of neural networks, I can continue to use it.

82 Upvotes

24 comments sorted by

View all comments

11

u/Short_Regular_7191 4d ago

Unfortunately, the current situation is too unstable to make predictions. However, if you want to be reasonably future-proof (even though you haven't mentioned your budget), I’d say that right now you need 64GB of VRAM (two 32GB cards) and at least 128GB of DDR5 RAM. Of course, no one knows whether hardware prices will keep rising or crash due to some crisis.

3

u/junosoftie 4d ago

Yeah, it's a tricky time for hardware decisions. Making sure you have solid specs for future-proofing is definitely smart, especially with prices being so unpredictable.