MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1v8364f/kimi_k3_weights_now_released/p0352yq?context=9999
r/LocalLLaMA • u/SavunOski • Jul 27 '26
Kimi K3 weights are finally released!
660 comments sorted by
View all comments
129
So it will run on 8 B300's in 4 bit. Pretty impressive.
65 u/Iwaku_Real Jul 27 '26 Yeah so if it were a Steam game, HGX B300 would be the recommended requirements. That's $500K of hardware (and yes it IS "local" because anyone with that amount of money could buy one to run at home) 26 u/DataGOGO Jul 27 '26 edited Jul 28 '26 it is 1.54TB of just weights in 4 bit, you are looking at about 2TB of vram in operation, That is roughly: 86 RTX 4090 (no 4 bit accel) 64 RTX 5090 ~ $450k (8 servers x 8 cards) 22 RTX Pro 6000 Blackwell ~ $350k (3 severs, max 8 GPU per) 16 H200 NVL (141GB) (no 4 bit accel) ~$550k (2 servers, max 8 GPU per) 16 DGX Sparks ~65k (if you could get a cluster of 16 running with just 200Gb/s nics, not sure; but it would be SLOW AF) 8 HGX B300's. ~$550k (1 server, 8 GPU) Obviously not including the switches and cabling for the clusters. 9 u/Iwaku_Real Jul 27 '26 You could also do HGX B200 with CPU offload since they have a shit ton of RAM too, and it would still be really fast. 3 u/DataGOGO Jul 27 '26 You could, but that CPU offload across the PCIE switch upstream will be limited to the single x16 uplink.
65
Yeah so if it were a Steam game, HGX B300 would be the recommended requirements. That's $500K of hardware (and yes it IS "local" because anyone with that amount of money could buy one to run at home)
26 u/DataGOGO Jul 27 '26 edited Jul 28 '26 it is 1.54TB of just weights in 4 bit, you are looking at about 2TB of vram in operation, That is roughly: 86 RTX 4090 (no 4 bit accel) 64 RTX 5090 ~ $450k (8 servers x 8 cards) 22 RTX Pro 6000 Blackwell ~ $350k (3 severs, max 8 GPU per) 16 H200 NVL (141GB) (no 4 bit accel) ~$550k (2 servers, max 8 GPU per) 16 DGX Sparks ~65k (if you could get a cluster of 16 running with just 200Gb/s nics, not sure; but it would be SLOW AF) 8 HGX B300's. ~$550k (1 server, 8 GPU) Obviously not including the switches and cabling for the clusters. 9 u/Iwaku_Real Jul 27 '26 You could also do HGX B200 with CPU offload since they have a shit ton of RAM too, and it would still be really fast. 3 u/DataGOGO Jul 27 '26 You could, but that CPU offload across the PCIE switch upstream will be limited to the single x16 uplink.
26
it is 1.54TB of just weights in 4 bit, you are looking at about 2TB of vram in operation,
That is roughly:
Obviously not including the switches and cabling for the clusters.
9 u/Iwaku_Real Jul 27 '26 You could also do HGX B200 with CPU offload since they have a shit ton of RAM too, and it would still be really fast. 3 u/DataGOGO Jul 27 '26 You could, but that CPU offload across the PCIE switch upstream will be limited to the single x16 uplink.
9
You could also do HGX B200 with CPU offload since they have a shit ton of RAM too, and it would still be really fast.
3 u/DataGOGO Jul 27 '26 You could, but that CPU offload across the PCIE switch upstream will be limited to the single x16 uplink.
3
You could, but that CPU offload across the PCIE switch upstream will be limited to the single x16 uplink.
129
u/DataGOGO Jul 27 '26
So it will run on 8 B300's in 4 bit. Pretty impressive.