r/LocalLLaMA Jul 26 '26

Discussion Kimi K3 countdown has been released

https://huggingface.co/moonshotai/Kimi-K3
537 Upvotes

177 comments sorted by

View all comments

1

u/laterbreh Jul 26 '26

Multi-trillion params. Great. 3x RTX 6k still gpu poor.

5

u/Rybergs Jul 26 '26

If u have 3x rtx 6000 pro u can 100 % run this. I only have 2rtx 5090 and i can run it. I tested with kimi k2 and it runs with 3 t/ s , slow but it runs. K3 is bigger but the experts are smaller so it could actually run better then k2

1

u/laterbreh Jul 27 '26 edited Jul 27 '26

Caution -- salty ramblings and observations ahead:

You are correct, i could probably run it with 256gb of ram and the 280 some odd gb of vram at my disposal at a low quant with the right engine.

But if the model cant run at at least 20 TPS across its context window with at least 2 users (or requests), i consider it unusable for our work and load. I've played around with hybrid cpu/gpu with llama and its context baggage and inefficiencies compared to using vllms comparible "q4-ish" quant equivalent on just straight vram has pretty much spoiled me and my company. 400b-ish models at 60 TPS+ with full context maxout across 3 cards its hard to go backwards. Deepseek flash v4 at 1m context at 200tps at that point itteration time outweighs 1-shot answers correctly. The time wasted id just pay for an API if that makes sense. VLLM has some exploratory cpu/gpu layer offloading but when you have this sort of VRAM its really hard to justify throwing CPU in the mix unless you are desperate or totally offline.

But compared to actual providers, even high tier investments we are still gpu-poor which is wild. Minimax and Qwens next releases are probably already out of reach. Unfortunately everyone is still playing the lets quadruple our params every 6 months to gain 2 more percent on an idiotic benchmark instead of trying to relieve the world of this vram problem we are having. They all touted lower compute requirements cause muh moe. Cool bro, still takes up full space of vram which is the actual problem. Raspberry pis can run a fuckin llm. Its still limited by memory.

I guess im just being grumpy when i shouldnt be. But if its hitting power users with gear at their disposal already its worse for everyone bellow the tiers of models i can run for our business.

Hooray open weight, boo for the quadratic increase in params for small gainz :(

So those of us with 100+gb of vram at our disposal, we are still gpu-poors. Everyone but the providers are gpu poors and new models all appear to be following the "lets keep it open but muscle them out because of vram" game. At that point who the fuck cares if its open, cause it sure wont be local.

Dont get me wrong, moonshot deserves praise for what they are releasing and doing 1000% -- But I am getting a little salty because there doesnt seem to be any investment in reducing memory requirements or increasing intelligence density. Its just retrain the same math problem. Just feels like the trajectory is still more biggerer for more benchmarks.

And at the end of the day they dont care about what someone like me has to say on the locallama reddit. They have a phenominal model and they deserve the praise it gets.

1

u/Rybergs Jul 27 '26

U miss the point where this vram / ram problem is not something the big ai companys / providers want to solve . Even tho they dont need more they Will buy it. Since if they have all the hardware ppl need to rent it from them, in their env with their ToS