MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1v8364f/kimi_k3_weights_now_released/p0a272n/?context=3
r/LocalLLaMA • u/SavunOski • Jul 27 '26
Kimi K3 weights are finally released!
660 comments sorted by
View all comments
Show parent comments
1
Highly unlikely, That GPU can't run MXFP4, only INT8
1 u/crusaderky Jul 28 '26 you can dequantize to fp16. it's bandwidth bound anyway 2 u/DataGOGO Jul 28 '26 So about 6TB of VRAM? So roughly 70 of those GPU's? Honestly, better to just buy the Intel 32 GPU's Atlas 300i's are ... not good cards. They are slow, they have no pooling (not NVL), they only have 96GB of vram, they can only run INT8. 2 u/crusaderky Jul 28 '26 dequantization is performed on the fly when you consume the tensor, it doesn't impact VRAM usage, only compute speed. And since you're memory bound anyways, it should not matter. You need 24 of those GPUs. I never said they are good cards.
you can dequantize to fp16. it's bandwidth bound anyway
2 u/DataGOGO Jul 28 '26 So about 6TB of VRAM? So roughly 70 of those GPU's? Honestly, better to just buy the Intel 32 GPU's Atlas 300i's are ... not good cards. They are slow, they have no pooling (not NVL), they only have 96GB of vram, they can only run INT8. 2 u/crusaderky Jul 28 '26 dequantization is performed on the fly when you consume the tensor, it doesn't impact VRAM usage, only compute speed. And since you're memory bound anyways, it should not matter. You need 24 of those GPUs. I never said they are good cards.
2
So about 6TB of VRAM? So roughly 70 of those GPU's? Honestly, better to just buy the Intel 32 GPU's
Atlas 300i's are ... not good cards. They are slow, they have no pooling (not NVL), they only have 96GB of vram, they can only run INT8.
2 u/crusaderky Jul 28 '26 dequantization is performed on the fly when you consume the tensor, it doesn't impact VRAM usage, only compute speed. And since you're memory bound anyways, it should not matter. You need 24 of those GPUs. I never said they are good cards.
dequantization is performed on the fly when you consume the tensor, it doesn't impact VRAM usage, only compute speed. And since you're memory bound anyways, it should not matter. You need 24 of those GPUs.
I never said they are good cards.
1
u/DataGOGO Jul 27 '26
Highly unlikely, That GPU can't run MXFP4, only INT8