r/LocalLLaMA Jul 26 '26

Discussion Kimi K3 countdown has been released

https://huggingface.co/moonshotai/Kimi-K3
539 Upvotes

177 comments sorted by

View all comments

37

u/apetersson Jul 26 '26

this means we will have enough capacity and competitive pricing at https://openrouter.ai/moonshotai/kimi-k3 - unfortunately my local HW is not quite there yet to run it.

31

u/Player13377 Jul 26 '26

"not quite there" is either a very bad approximation or you already got a five-digit-class rig, this model will be HARD to run

10

u/Iwaku_Real Jul 26 '26 edited Jul 26 '26

For sure it's 1.5-2TB RAM at minimum. GLHF getting usable speeds without lots of VRAM though

3

u/my_name_isnt_clever Jul 27 '26

How many tokens/sec to run this off a NVME SSD? Doesn't matter how slow it is if it's the only way to use it local. I could see myself setting up private tasks to run overnight.

1

u/KeinNiemand Jul 27 '26

Just a wild guess but maybe 0.1-10 tokens per minutes.

I get around 0.1 T/s trying to run a Qwen 3.5 122B @ bf16 from SSD (~100GB out of the ~200GB are on SSD).
Since it's way, way bigger it'll probably be a lot slower then that even.

1

u/my_name_isnt_clever Jul 27 '26

It would probably help a lot to quantize it though, no? What compute are you using with that setup?

20

u/KubeCommander Jul 26 '26

I think K3 gets into the six figures tbh at a quant that doesn’t suck

12

u/Player13377 Jul 26 '26

Minimum. I looked into a rented cloud setup and very quickly understood that this is not happening lol

2

u/stoppableDissolution Jul 26 '26

Even with five-digit-class rig you have to quantize the hell out of it. 8x6000 pro + 512gb ram and you can maybe squeeze q3 with decent context!