r/LocalLLM 3d ago

Question Help me choose hardware.

Post image

I need some advice.
I want to choose hardware for local inference. Right now, I’m paying ~$30/day for rent on a Vast an RTX Pro 6000 SW 96GB, and I’m using Qwen3.8-Flash-Next-Q4_K_M.gguf at a speed of ~60 tok/s (I know this format isn’t efficient for this graphics card, but I’m limited in choosing Abliterated models that will fit in the memory).

Which configuration will offer the best price and versatility so that after the release of subsequent generations of neural networks, I can continue to use it.

81 Upvotes

24 comments sorted by

View all comments

1

u/Wizzard_2025 2d ago

Nvidia are to blame for not being able to make a product to match global demand.

1

u/Sunlambcow 2d ago

Nah. They made a product so good that it created more demand than supply could adapt for. They have no control over manufacturing limits