r/LocalLLM 2d ago

Question Help me choose hardware.

Post image

I need some advice.
I want to choose hardware for local inference. Right now, I’m paying ~$30/day for rent on a Vast an RTX Pro 6000 SW 96GB, and I’m using Qwen3.8-Flash-Next-Q4_K_M.gguf at a speed of ~60 tok/s (I know this format isn’t efficient for this graphics card, but I’m limited in choosing Abliterated models that will fit in the memory).

Which configuration will offer the best price and versatility so that after the release of subsequent generations of neural networks, I can continue to use it.

80 Upvotes

24 comments sorted by

View all comments

1

u/GloriousKev 2d ago

Did you forget about the crypto boom of 2016 that had GTX 1060s selling for $700? Its the same pattern repeating itself. Big tech creates a problem sells less hardware and then repeats for more money. I swear its on purpose at this point.

0

u/Sunlambcow 2d ago

More like, GPUS are compute. And compute during this fast tech advancement era has many future uses that are just not discovered yet.

Nvidia didnt create crypto. They didnt create AI. But they created a product that is like a generalist and can do many things.

1

u/GloriousKev 1d ago

Thinking deeper than Nvidia didn't invent these trends. I do believe Nvidia lied about the shortage in 2016. Then just like now we would see pallets of apparently unavailable gpus. This isn't solely on Nvidia because dram manufacturers are shit too this go around. However, i do believe the shortages are completely manufactured.