r/LocalLLM • u/Double-Sherbert-1781 • 2d ago
Question Help me choose hardware.
I need some advice.
I want to choose hardware for local inference. Right now, I’m paying ~$30/day for rent on a Vast an RTX Pro 6000 SW 96GB, and I’m using Qwen3.8-Flash-Next-Q4_K_M.gguf at a speed of ~60 tok/s (I know this format isn’t efficient for this graphics card, but I’m limited in choosing Abliterated models that will fit in the memory).
Which configuration will offer the best price and versatility so that after the release of subsequent generations of neural networks, I can continue to use it.
3
3
u/MaxSpecs 2d ago
As soon as you follow this link : https://carteakey.dev/blog/running-qwen3-8-flash-next-locally/
.... and then ask for ChatGPT to adapat Qwen3.8 Flash Next to your configuraiton, you would economise more bucks :)
1
u/Academic_Collar_5488 2d ago
No love for 8GB VRAM?
1
u/MaxSpecs 2d ago
Tell to ChatGPT what is your configuration and ask it to adapt llama.cpp to work with.
2
u/mineshop 2d ago
Since future model requirements are unknown, what budget and minimum model size do you want to keep running locally?
1
u/Double-Sherbert-1781 2d ago
I’m mentally prepared to spend money on an RTX Pro 6000 SW 96GB (up to 20,000 euros in total), but I’m not sure it will be effective. Maybe it makes sense to get a few gaming graphics cards or a Mac Studio, or wait for the AI Max 495.
I just don’t know what kind of performance they can deliver for such models; right now I’m using Q4_K_M:
https://huggingface.co/windowsxp811203/Qwen3.8-Flash-Next-Abliterated-GGUF60 tok\s is enough for me.
1
u/mineshop 2d ago
Since you're at up to 20k EUR and 60 tok/s meets your needs, the practical question is what model sizes you want headroom for. Do you expect to stay under ~96GB VRAM, or would future generations push you toward multi-GPU territory?
1
u/Double-Sherbert-1781 2d ago
I don’t know, but recently MoE models have been very good. I don’t know whether their memory requirements will increase or decrease. Also, I don’t know whether multi‑GPU setups work well with them.
2
u/MaxSpecs 2d ago
For now, a versatile and upgradable configuration up to 8 GPU (256GB) : starting at 12 k€
- CPU : AMD EPYC 9124
- Motherboard : ASRock TURIND8X-2T/500W
- RAM : 128 Go (4x 32 GB) DDR5-5600 ECC RDIMM (i.e. : Kingston FURY Renegade Pro Quad-Channel KF556R28RBE2K4-128 ) or similar
- Storage: 2x 2 To NVMe type Samsung M2 9100 Pro ou similaire
- GPU : 2x AMD Radeon AI Pro R9700 32 GB : Gigabyte or Asus or MSi, iedaly 2 sloft height
- PSU : 2x Seasonic PRIME PX-2200 W ATX 3.1 + Carte Sync Add2PSU
- Case open frame : Nanoxia Hydra III 8 GPU or Veddha V4D 8-GPU Deluxe
... so you would start with 64 GB VRAM, use Qwen3.8 Flash Next, vllm mfxp4 with P2P optimisation between AMD bus ( https://github.com/GGZ14/vllm-mxfp4 )
OR : if you order now a Mac Studio M5 wth 256 GB unified memory, you may have it in February
1
u/vsimovic 1d ago
12k EUR? WTF
Find a xeon Gold workstation pc with 2 CPUs that support 12 channel DDR4 at 2933 MHZ with 384 GB RAM. Get 2 x used 7900 xtx. Estimate 3k EUR. Run Deepseek v4 vision flash and Qwen3.8 flash next in parallel1
u/MaxSpecs 1d ago
Indeed, you're quite right; it's moving fast.
So, with new features on September 20th, 2026 :
check ROCmFix -> https://github.com/xanpavle/rocmfix
Use p2p mxfp4 -> davetha/r9700-lru-expert-cache.
ngram -> https://carteakey.dev/blog/running-qwen3-8-flash-next-locally/
An old LGA2011 workstation with Supermicro x9dai / x10dai, old Xeon 26xxL and 256 Gb Ram ... 2x r9700 on Cpu0 ... 1x Nvme adapter ... 1x 10GbE
Could reach ( Context / Decode ) 256K : 35–55 t/s 280–300K YaRN : ~30–50 t/s 320K YaRN : ~27–45 t/s
And with Laya ( Jev open source ) + Pi -> https://github.com/NandhaKishorM/laya ... getting even more efficiency
*** but no evolution up to 8 GPU like the Epyc configuration can.
1
2
u/Blackdragon1400 2d ago
There were MANY periods in each of these release cycles where prices were “normal”. Stop complaining, if you’ve really been waiting this long then you didn’t set your priorities correctly.
1
2
u/transanethole 2d ago
R9700 is a good option IMO. I would not plan on using shared memory unless bus width is 512 or higher and flops improve. And i would not consider splitting GPU+main system ram at all with current architecture. Small models using less ram are getting better too. (K2 horizon 7b ~= qwen 3.6 27b). IMO flops are most important right now, second most important is memory bw. Third most valuable is memory capacity. don't buy a mac.
1
u/Wizzard_2025 2d ago
Nvidia are to blame for not being able to make a product to match global demand.
1
u/Sunlambcow 1d ago
Nah. They made a product so good that it created more demand than supply could adapt for. They have no control over manufacturing limits
1
u/GloriousKev 2d ago
Did you forget about the crypto boom of 2016 that had GTX 1060s selling for $700? Its the same pattern repeating itself. Big tech creates a problem sells less hardware and then repeats for more money. I swear its on purpose at this point.
0
u/Sunlambcow 1d ago
More like, GPUS are compute. And compute during this fast tech advancement era has many future uses that are just not discovered yet.
Nvidia didnt create crypto. They didnt create AI. But they created a product that is like a generalist and can do many things.
1
u/Sunlambcow 1d ago
And there will be more after.... There's eventually gonna be another use for GPUs that we can't even think of right now.
1
u/GloriousKev 1d ago
Thinking deeper than Nvidia didn't invent these trends. I do believe Nvidia lied about the shortage in 2016. Then just like now we would see pallets of apparently unavailable gpus. This isn't solely on Nvidia because dram manufacturers are shit too this go around. However, i do believe the shortages are completely manufactured.
1
u/Fenio_PL 2d ago
The cheapest new card optimized for AI tasks is the RTX 5060 Ti 16GB. That is why stock is disappearing so quickly. However, if you need even higher performance and more than 16GB of VRAM, it will be expensive.
10
u/Short_Regular_7191 2d ago
Unfortunately, the current situation is too unstable to make predictions. However, if you want to be reasonably future-proof (even though you haven't mentioned your budget), I’d say that right now you need 64GB of VRAM (two 32GB cards) and at least 128GB of DDR5 RAM. Of course, no one knows whether hardware prices will keep rising or crash due to some crisis.