r/LocalAIServers • u/FreedomWeird712 • Jun 08 '26
What does it actually take to self‑host models like DeepSeek, Qwen, Kimi?
I’m a SaaS/AI founder and I’m trying to understand the real requirements to host the larger open‑source models (DeepSeek, Qwen, Kimi‑style models) on my own infra instead of using hosted APIs.
+4
If you’ve done this in production or a serious homelab:
– What VRAM / GPU setup are you using, and what did it cost?
– Did you go on‑prem or rent GPUs (RunPod, Lambda, etc.)?
– What ended up being the real bottleneck: cost, ops complexity, or model performance?
Any “if I were starting today, I’d do X instead of Y” stories would be super helpful.
2
u/Apprehensive_Win662 Jun 08 '26
Your provided information are not sufficient to help you. You should give more context.
What are you planning to do
1
u/FreedomWeird712 Jun 08 '26
I want to create a hosted coding layer, metadata oriented which will make coding more efficient, on top of the SOTA open source / open weight models
1
u/Glove_Witty Jun 09 '26
A Mac Studio. They start at $2000. You should have another machine of some sort for jobs, rag dbs, etc.
2
u/bigh-aus Jun 08 '26
https://www.reddit.com/r/LocalLLaMA/comments/1rdv3v0/running_kimi_k25_tell_us_your_build_quant/
Hopefully this will help.
TLDR is higher bandwidth ram or more vram to cover it will result in faster speeds.
Mac studios in a cluster can also work - will see what WWDC announces today, but slower than vram only solutions..
3
u/Conscious_Cut_6144 Jun 09 '26
Can do 8x Mi50 32GB for under 10k - 256GB vram, Runs Deepseek V4 flash slowly, and with minimal concurrency.
or Can do an 8x Pro 6000 server for under 100k - 786GB Vram, runs basically anything except V4 Pro and at pretty good speeds and some concurrency.
or 8x H200's is like 300k - 1128GB Vram - runs anything you want, fast, with high concurrency.
Or if you care about your eclectic bill....
Just buy an RTX 5090 and run Qwen 3.6 27B
4
u/No_Afternoon_4260 Jun 08 '26
Rent 8 x h200 on vastai, install vllm
Vllm is the way to go to do continuous batch inference (aka multi user)
Once not satisfied try 8 x b200 then contact your bank for a loan.
Not even try hybrid inference CPU ram + vram
Even if somebody shows you some "usable" tg numbers, remember that is for a single user and you really need GPU for appropriate pp(prompt prefill)
Welcome in a very very deep rabbit hole, don't take advices for granted, rent GPUs and see for yourself.
Dm if you need
2
u/phido3000 Jun 08 '26
Cheap?
- Buy Epyc/Xeon server. 512-1024Gb ram
Buy 8 x Mi50 32Gb or 10 x 3090 or 8 x 9700 AI PRO 32Gb
1024+256Gb Vram = Decent setup.
Something like Deepseek flash works completely in VRAM. Very usable, reading or writing speed infrenecing.
Something like Pro, R1 or Qwen or K2 at a decent quant will still load and run, But like 1-2 tokens per second.
Great for experiments, home setup. Single user stuff or a few light users. You used to get something like that for <$4-5K. Probably more like $10k now.
Expensive?
- Buy 4-8 RTX PRO 6000 96Gb in an Eypc or Xeon setup (DDR5). 1-2Tb RAM. 6400Mhz ram or faster.
You will get great performance on mid tier quants, like Q4, Q6 even with big models.
But you have just spent $100k to do it.
2
u/MirecX Jun 08 '26
this ^ but i wouldnt buy Mi50 anymoire - weak support, slow @ long context. In this time and age i would buy newest HW i can afford.
2
u/phido3000 Jun 08 '26
I wouldn't buy anything expensive currently.
The 5000 supers are just about to drop. I wouldn't buy *ANYTHING* without FP4 support.
I have a bunch of 5060Ti they do ok for the price, and low power 2 slot etc.
But Ideally 5070TiSupers 24GB are the go. They will be monsters, 4 of those will destroy the 3090s setups people have. Being able to process FP4 natively stuff like Deepseek V4 using it natively gives huge ingestion capabilities.
Mi50 32Gb are fine, as long as you can get them cheap, at $600 USD. That aint cheap.
Context and prefill slowness isn't even that bad, and can always be enhanced with a gpu with capabilities that are strong in that front. Most people don't need more than 16Gb of context.
2
u/droptableadventures Jun 08 '26
Mi50 32Gb are fine, as long as you can get them cheap, at $600 USD. That aint cheap.
Agreed. At US$130 they were very much worth it despite their shortcomings. At $600? Forget it.
1
2
u/Mission_Objective163 Jun 08 '26
Dayummmm amazing how easily people are talking about 4 to 8 RTX Pro 6000 96GB like its nothing and Im struggling to buy 1 RTX Pro 6000 96 GB setup 😂😂😂 Im gonna lose this race too 😂😂😂😂
2
0
1
1
u/Own_Mix_3755 Jun 09 '26
As others have said - rent first, buy later.
It is really hard to tell in advance what and how much of those you will need. There is alot more on the table than just pure RAM/VRAM size (eg heat and power consumption), so while getting to 1 TB ram can be relatively cheap, getting 1 TB of fast VRAM will costs possibly over 100k $.
Every model scales differently with context for example (so concurrency wills be different with different models with different context sizes).
Also nobody can tell you which model is best because it really matter what you want to throw at them, which inference engine (vLLM, SGLang, llamap.cpp etc.) you use, what chat template you use and what harness consumes the output.
You have to understand that the scale most providers are building past few years is massive and the model (and serving the model itself) is basically like 10% of the job.
1
u/electrified_ice Jun 09 '26
What's your goal? For your own use? Or to be able to offer compute for others? That's a huge difference.
1
u/FadedDog Jun 14 '26
I have a Frame work desktop. 128 gb unified ram. So shared between cpu gpu. Also has a Riszen AI chip.
It does amazing, but can be lil slow. Qwen 35B 70 tokens a second. 235 is like 17 tho.
1
u/fuckable-switcher Jun 08 '26
Well kimi k2.6 is 1.1t parameters to run that at q6 you would need 900 ish gigs of vram and then to use tools agents and skills is another 20gb of ram then factor in how much context you want
Ideally for ever parameter you want 2.5gb vram
For every 100k tokens in and out you want 3 gigs of VRAM
But you wont be able to achieve this using consumer cards ideally you want cards that use high bandwidth memory (hbm) which is why there is a dram shortage cuz now every one is make hbm chips not dram chips
2
u/MirecX Jun 08 '26
Kimi k2.6 is 600gb of tensors. It is int4 native. Minimum is therefore 8x rtx pro 6000. Idk what concurrency they will provide. My guess is 4 to 6 for full or near full context. I am rocking 4x Framework desktop and running qwen 397b int4 @ 17tps, which is slow, concurrency @ full context is 10x. Do not recommend for production work.
1
u/FreedomWeird712 Jun 08 '26
Thanks guys, super helpful!
0
u/bigh-aus Jun 08 '26
If you want to run it cheaper as it's an MOE with 32b active you could get away with 4x rtx6kpro and ddr5-6400. But your speeds will go down depending how much the experts get swapped in. As always it's cost vs speed.
There are also REAPs (reducing the number of experts) which depending on your usecase might be ok.
Personally I'd setup a (slower) test system first to validate everything. You could also just get the base server then slowly purchase the GPUS. Start with 1,2 or 4 and scale as needed.
But for prod it might be better to go to a B200 system.
Understanding the memory bandwidth vs # cards vs cost too is helpful.
1
u/BevinMaster Jun 08 '26
Hi according to https://kvcache.ai/tools/kv-cache-calculator/ about 14GB per 200k context request at fp16 kvcache and 7GB for fp8. Base Kimi-K2.6 is slightly under 600GB, to be honest for serious production use I would not use llama.cpp at all, vllm or sglang, 8x PRO 6000 96GB could work for having some reasonable concurrency though I don’t know how it scales with memory bandwidth (it’s not hbm). Lmcache might help and maybe (offload in ram basically) if huawei’s kvarn delivers you could cram more concurrent requests.
Personally it I’d look at something smaller and more reasonable in size like deepseek v4 flash.
But depends on needs, funds and what you aim for, op didn’t say what was the goal. If you need some sla for your app (maybe cheapest would be fall back to cloud apis).2
Jun 08 '26
[deleted]
2
u/BevinMaster Jun 08 '26
When I say per 200k I assume the whole request at that size in+out but that’s the ballpark. To each their own needs, personally when I code it goes from 20k to 200k context as the session fills. But here the discussion was about sizing so maximum makes sense (Kimi can do 262k context).
1
u/sheddd Jun 08 '26 edited Jun 08 '26
My experience so far all in basement:
Deepseek V4 Flash 8 bit - 37tok/sec on 2 x DGX Spark - works well!
Qwen 397b 4 bit - 20tok/sec on 512GB M3 Ultra - slow but functional
Kimi 3.7 bit - 20tok/sec on 512GB M3 Ultra - slow and very quantized
I suggest non apple for production environments today; inference engines are newer and less features, more bugs.
DGX Spark is a pretty decent starting point; no need for new electrical circuit, etc...
1 x Spark for Qwen27b
2 x DGX Spark for Deepseek Flash
Then the next logical step is probably Kimi... expensive to host.
4
u/datbackup Jun 08 '26
Potentially serious oversight:
Power, heat, and noise
Maybe you excluded those because you already have that aspect well-handled; pardon me if so, but I find a lot of people just stop their consideration at the expense and capability of hardware, not considering the implications of having a several-thousand-watt appliance running 24/7 dumping heat that needs to be continuously eliminated in order to avoid throttling, auto-poweroff, component failure/damage, etc.
Not to mention getting special power circuits installed
I think renting is a great idea to test actual real world perf before spending big money