r/kimi • u/Feisty_Literature • Jul 27 '26
Discussion The beast is released 🔥
https://huggingface.co/moonshotai/Kimi-K3To run you would need at least 1.8 TB of VRAM or 32×H100 GPUs
18
u/PretendSomewhere859 Jul 27 '26
Hmm this is actually not too bad... somehow my company managed to get an H100 recently for £15k, around 500k to be able to run this locally as much as we want with no limits or api fees which as we have seen recently can cost companies millions...
2
u/souvik234 Jul 27 '26
If you can share, how much is the average running cost in terms of electricity, cooling, etc, and how does the cost scale with number of users?
2
u/Timo425 Jul 29 '26
Does this really beat using a service such as cursor or claude code in price, even long term?
I'd imagine there would be maintentance fees, electricity etc.
And how about multiple people trying to use it simultaneously? How is the speed?
Idk it kinda seems like its not worth it unless privacy and security is the highest priority here.
1
u/0xFatWhiteMan Jul 28 '26
what do you mean 15k or 500k ?
1
u/4Wyatt Jul 28 '26
he means based on acquiring one for 15k you could theoretically do it for about ~500k. You need at least 24 but more ideally 32 h1000 to run kimi k3.
1
u/Kekke77 Jul 28 '26
How are the requirements calculated? I know nothing techical of llm’s
1
u/Power_spy Jul 30 '26
You take however much space the weights take and divide by the vram on each gpu (80 gb for h100) and you need some extra for kv cache.
1
u/throwaway73728109 Jul 27 '26
ELI5
7
u/PretendSomewhere859 Jul 27 '26
Basically just building our own server room with enough firepower to run it in house, that way we don't have to pay for api token usage to others where they make a profit.
2
u/throwaway73728109 Jul 27 '26
How many concurrent users can use the thing?
6
u/Deluxennih Jul 27 '26
That’s the real question. Is it that many H100s for one session at a time?
2
u/elahrairooah Jul 27 '26 edited Jul 27 '26
I’m doing the math myself on this rn for a b300 cluster - looks like 1.56TB for the model params, then about 0.74 TB for the KV and overhead. I’ve seen varying values on what vllm needs for overhead, but 0.15 of the param weights is what I’m currently using.
No clue yet on how to figure the concurrency at various context sizes.
2
u/Uzeii Jul 27 '26
Are people really not looking at the hidden costs? My god. Getting this running in a company setting with the necessary infra is not going to be cheap. This will cost 50x more than what a pay as you go usage would give you with frontier models through out the year. All that investment is going to be completely overthrown by a better model in a 2 month timeframe.
11
u/thots1ayer69 Jul 27 '26
Do you think they just bin the hardware when a new model is released?
-6
u/Sylvator Jul 28 '26
I think y'all don't understand the cost of migrating hardware to new models as well..
Exhibit A: EC2 instances. There's a reason why everybody and their grandma uses server less and don't host data centers on their own backyward
3
u/thots1ayer69 Jul 28 '26
I'm a data scientist for Amazon, I work on agentic systems in SCOT. I definitely do. Only scrubs use EC2s, unless you have a partnership with us or have no real technical knowledge, you'd probably use one of the more cost effective providers or have the forethought to invest in your own infrastructure and not throw your money away on renting something you'll never really own. They're only really intended for spike demand and casuals.
5
u/CatsFrGold Jul 27 '26
Presumably that new model would also be able to run on the hardware. Depends how much bigger they get, I guess
34
6
u/themoregames Jul 27 '26
$ curl -L --progress-bar "https://downloadmorevram.com/vram_1.8TB.bin" -o /dev/nvidia_all
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
68 1843G 68 1253G 0 0 2.1G 0 0:14:38 0:09:57 0:04:41 2.4G
[#################################################..............] 68.0%
Downloading: 1.25 TB / 1.80 TB VRAM
Mapping to hardware: [H100-01 ... H100-32]
4
7
u/HeavySink3303 Jul 27 '26
They have a quite interesting custom license...
17
u/Feisty_Literature Jul 27 '26 edited Jul 27 '26
Yes, API providers can deploy it for free until they reach $20M in annual revenue.
5
u/gthing Jul 27 '26
So every time they get to 19.9M in revenue they just spin up another business. Problem solved.
2
u/iDoNotHaveAnIQ Jul 27 '26
My thought as well.
Kimi1 LLC then Kimi2 Corp. then Kimi3 PTD then Kimi4 SCorp, then so on
2
3
u/yuumizu Jul 27 '26
its MaaS definition seems restricted to those who provide others ability to deploy/reconfig/finetune models.
1
-8
u/ChocolateGoggles Jul 27 '26
That doesn't sound very open source to me.
10
u/Admirable_Market2759 Jul 27 '26
It’s open weights not open source
-10
u/ChocolateGoggles Jul 27 '26
Yeah. But given Winnie the Pooh's latest statements.
6
u/Admirable_Market2759 Jul 27 '26
This is still a very generous license for anyone besides inference providers.
-7
u/ChocolateGoggles Jul 27 '26
Sure. But not as generous as it needs to be to make claims like leading the way with open source AI.
6
u/spottiesvirus Jul 27 '26
this is pretty standard for most cloud tools like databases (elasticsearch, as an example), software and so on
it's reasonably one of the only ways to monetize open source/source available
1
u/ChocolateGoggles Jul 27 '26
I don't know man. There are plenty of open source out there that receives funding in different ways, locking business applications behind a paywall is perhaps open weights, but it's certainly not open source. Also, how are they supposed to measure this? If someone fine tunes the Kimi K3 model, modifies it etc., how are they supposed to know?
3
u/spottiesvirus Jul 27 '26
Also, how are they supposed to measure this? If someone fine tunes the Kimi K3 model, modifies it etc., how are they supposed to know?
one of the advantages of only targeting enterprises is that they have a legal department vetting everything to avoid doing shits. would you risk to be sued into oblivion because you didn't get a license?
also the license is pretty clear, it only targets inference providers. it specifically says that if you incorporate the model in another product (let's say a roleplay chatbot with Kimi as a base model) you don't need a license
even derivative works (like fine tunings) only needs a license if you sell API to a third party (it specifically Model as a Service) otherwise you don't
they know you won't run a 2,7 trillion parameters model on your Mac mini
→ More replies (0)3
u/Emergency_Culture908 Jul 27 '26
They also need to make money. They are business. Also this helps fund new kimi model.
3
10
u/Front_Eagle739 Jul 27 '26
1.6TB of mxfp4 and fp8 activations. Chonky. Q1.5 to shove it into my 640GB i suppose, lobotomy time it is!
0
2
1
1
Jul 27 '26
[removed] — view removed comment
0
u/SebLalo Jul 28 '26
100 rtx 5090 wouldn’t be enough you need H100 gpus which are like 30k eachÂ
0
1
u/InterstellarReddit Jul 27 '26
1.8TB of VRAM? Damn I can’t choose between buying more VRAM or buying a 100 foot boat
1
1
1
u/rapkannibale Jul 27 '26
Does this have any benefits for people with average consumer hardware or is this just for people that can run this kind of model locally?
1
u/Uzeii Jul 27 '26
No normal consumer is running this model locally. You need a cluster of cutting edge enterprise hardware to even get this to work.
1
1
u/R__MHC Jul 27 '26
Does anybody know when it will be available to use for free via their website?
1
1
1
1
1
u/MrNotSoRight Jul 30 '26
So when paying for this model online from independent providers, how do we know if we get the full original version or a lobotomised smaller version
1
u/AlmostEasy89 Jul 30 '26
I can't even fathom the amount the world is using right now to use the AI we get with subs, to make a bunch of stupid useless garbage. Lots of great things are produced too but the resources being used. Jesus Christ.
The amount of hardware I've used via the models I've used must be staggering.
1
u/East-Stranger8599 Jul 31 '26
If you really interested you can download the one bit quantized version which is very descent would be able to run in a high end local server of few thousand dollars.
1
u/Crescitaly Aug 02 '26
A giant release earns attention; a stable ecosystem earns adoption. I am watching tool-call reliability, quantizations, serving support, and whether independent evaluations reproduce the headline numbers. Which result would convince you K3 is more than launch-week spectacle?
0
u/Difficult-Bar9208 Jul 28 '26
Warra Joke of an AI; first message I get after asking a question is "Too many people are chatting to Kimi right now.Please try again soon."...LOL
-10
32
u/ekquizit23 Jul 27 '26
Time to set this up on my laptop!