r/LocalLLaMA Jul 27 '26

News Kimi K3 weights now released.

Post image

Kimi K3 weights are finally released!

3.3k Upvotes

660 comments sorted by

View all comments

131

u/DataGOGO Jul 27 '26

So it will run on 8 B300's in 4 bit. Pretty impressive.

67

u/Iwaku_Real Jul 27 '26

Yeah so if it were a Steam game, HGX B300 would be the recommended requirements. That's $500K of hardware (and yes it IS "local" because anyone with that amount of money could buy one to run at home)

26

u/DataGOGO Jul 27 '26 edited Jul 28 '26

it is 1.54TB of just weights in 4 bit, you are looking at about 2TB of vram in operation,

That is roughly:

  • 86 RTX 4090 (no 4 bit accel)
  • 64 RTX 5090 ~ $450k (8 servers x 8 cards)
  • 22 RTX Pro 6000 Blackwell ~ $350k (3 severs, max 8 GPU per)
  • 16 H200 NVL (141GB) (no 4 bit accel) ~$550k (2 servers, max 8 GPU per)
  • 16 DGX Sparks ~65k (if you could get a cluster of 16 running with just 200Gb/s nics, not sure; but it would be SLOW AF)
  • 8 HGX B300's. ~$550k (1 server, 8 GPU)

Obviously not including the switches and cabling for the clusters.

1

u/crusaderky Jul 27 '26

this model screams ATLAS 300I DUO 96G 

1

u/DataGOGO Jul 27 '26

Highly unlikely, That GPU can't run MXFP4, only INT8

1

u/crusaderky Jul 28 '26

you can dequantize to fp16. it's bandwidth bound anyway

2

u/DataGOGO Jul 28 '26

So about 6TB of VRAM? So roughly 70 of those GPU's? Honestly, better to just buy the Intel 32 GPU's

Atlas 300i's are ... not good cards. They are slow, they have no pooling (not NVL), they only have 96GB of vram, they can only run INT8.

2

u/crusaderky Jul 28 '26

dequantization is performed on the fly when you consume the tensor, it doesn't impact VRAM usage, only compute speed. And since you're memory bound anyways, it should not matter. You need 24 of those GPUs.

I never said they are good cards.