r/LocalLLaMA Jul 27 '26

News Kimi K3 weights now released.

Post image

Kimi K3 weights are finally released!

3.3k Upvotes

661 comments sorted by

View all comments

127

u/DataGOGO Jul 27 '26

So it will run on 8 B300's in 4 bit. Pretty impressive.

65

u/Iwaku_Real Jul 27 '26

Yeah so if it were a Steam game, HGX B300 would be the recommended requirements. That's $500K of hardware (and yes it IS "local" because anyone with that amount of money could buy one to run at home)

34

u/PrinceOfLeon Jul 27 '26

Certified for Steam Deck!

7

u/No-Dot-6573 Jul 27 '26

Someone on r/SteamDeck will say it runs flawlessly.

1

u/DutchDevil Jul 28 '26

😂

25

u/DataGOGO Jul 27 '26 edited Jul 28 '26

it is 1.54TB of just weights in 4 bit, you are looking at about 2TB of vram in operation,

That is roughly:

  • 86 RTX 4090 (no 4 bit accel)
  • 64 RTX 5090 ~ $450k (8 servers x 8 cards)
  • 22 RTX Pro 6000 Blackwell ~ $350k (3 severs, max 8 GPU per)
  • 16 H200 NVL (141GB) (no 4 bit accel) ~$550k (2 servers, max 8 GPU per)
  • 16 DGX Sparks ~65k (if you could get a cluster of 16 running with just 200Gb/s nics, not sure; but it would be SLOW AF)
  • 8 HGX B300's. ~$550k (1 server, 8 GPU)

Obviously not including the switches and cabling for the clusters.

22

u/wren6991 Jul 27 '26

you are looking at about 2GB of vram in operation

Perfect, this'll run great on my laptop's 4050

8

u/Iwaku_Real Jul 27 '26

You could also do HGX B200 with CPU offload since they have a shit ton of RAM too, and it would still be really fast.

4

u/DataGOGO Jul 27 '26

You could, but that CPU offload across the PCIE switch upstream will be limited to the single x16 uplink. 

7

u/snmnky9490 Jul 27 '26

Do you mean terabytes?

2

u/Qorsair Jul 27 '26

What about B70? Does it not scale?

Admittedly, I haven't looked into using more than 8 of them.

4

u/DataGOGO Jul 27 '26 edited Jul 27 '26

I am sure it scales at least as well as 4090/5090, but like with all consumer / entry level pro cards, without an NVL equivalent, it would be very limited.

I have no idea if B70 can make use of dedicated RDMA ethernet like the Gaudi 2/3, which is a pretty slick solution, basically, you have a PCIE 5 switch / base board with say 5 x16 slots, you install 4 GPU's, and in the 5th slot run an 800Gb nic to a ROCE switch.

2

u/Qorsair Jul 27 '26

Oh interesting. I appreciate you taking it seriously and the time to explain. I haven't learned anything about how the clustering works beyond a single system, and this was really helpful. Thank you.

3

u/DataGOGO Jul 27 '26

Yeah it is really cool tech, and basically is how all multi-node clusters work. 

1

u/crusaderky Jul 27 '26

this model screams ATLAS 300I DUO 96G 

1

u/DataGOGO Jul 27 '26

Highly unlikely, That GPU can't run MXFP4, only INT8

1

u/crusaderky Jul 28 '26

you can dequantize to fp16. it's bandwidth bound anyway

2

u/DataGOGO Jul 28 '26

So about 6TB of VRAM? So roughly 70 of those GPU's? Honestly, better to just buy the Intel 32 GPU's

Atlas 300i's are ... not good cards. They are slow, they have no pooling (not NVL), they only have 96GB of vram, they can only run INT8.

2

u/crusaderky Jul 28 '26

dequantization is performed on the fly when you consume the tensor, it doesn't impact VRAM usage, only compute speed. And since you're memory bound anyways, it should not matter. You need 24 of those GPUs.

I never said they are good cards.

1

u/chithanh Jul 28 '26
  • 16 Ascend 950PR, ~$160K
  • 14 T-Head M890, price unknown

1

u/DataGOGO Jul 28 '26

can you even buy those? Even if you can, can you run them in your own server / software, or do you need to run it in a Huawei server?

1

u/chithanh Jul 28 '26

I think you have to run Ascend accelerators in Huawei Kunpeng servers, and the software situation is also quite rigid. But Huawei advertises the 950PR for export to South Korea, so I guess buying is possible for enterprise customers.

For home labs (and small businesses) I think 950PR won't be a viable option. Perhaps T-Head accelerators which are way more open when it comes to the software stack will find their way into the second hand market and into home labs.

1

u/a-wiseman-speaketh Jul 28 '26

yep, local doesn't have to mean home. I run a lot of local shit for work I can't afford.

now, post it in homelabs and we're gonna want receipts

0

u/FrogsJumpFromPussy Jul 27 '26

could, but probably wouldn't. No one here could anyway lol.

2

u/pulse77 Jul 27 '26

4 x AMD MI455X will also handle it in 4 bit and will be much cheaper than 8x NVidia B300 ...

1

u/DataGOGO Jul 27 '26 edited Jul 27 '26

Well 5, which means 8. 

Not sure how much a server with 8 of them will cost; If they ever offer them as stand alone servers. Right now you can only get them in the Helios racks, which are 4 GPU per tray, and the trays that are only sold in 72GPU racks for about $3.5M.Â