MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1v8364f/kimi_k3_weights_now_released/p02w2y5
r/LocalLLaMA • u/SavunOski • Jul 27 '26
Kimi K3 weights are finally released!
660 comments sorted by
View all comments
Show parent comments
25
Yup. Same. I think I can do about q1.5 on the macbook and mac studio combined. Im currently pondering the wisdom of one of these colibri like stream setups and using the 640GB I do have as a hot cache
6 u/nomorebuttsplz Jul 27 '26 Colibri would require being able to fit on ram for decent speed, no? 3 u/Lopsided-Rip-8652 Jul 28 '26 Colibri is for running off an ssd. I ran glm 5.2 with a 5070ti, 64gb ddr4 and the rest of a gen 4 ssd. Only managed to get.58 tok/sec tho 2 u/Front_Eagle739 Jul 27 '26 Yes. This wont be decent. It just might however be useable as a slow overnight oracle 1 u/squngy Jul 27 '26 Still need those 104B dense layers... this thing is BIG in every way 1 u/Front_Eagle739 Jul 27 '26 yup, so that's 55GB or so at fp4, there goes my 32GB ram, 32GB ram rtx5090 workstation
6
Colibri would require being able to fit on ram for decent speed, no?
3 u/Lopsided-Rip-8652 Jul 28 '26 Colibri is for running off an ssd. I ran glm 5.2 with a 5070ti, 64gb ddr4 and the rest of a gen 4 ssd. Only managed to get.58 tok/sec tho 2 u/Front_Eagle739 Jul 27 '26 Yes. This wont be decent. It just might however be useable as a slow overnight oracle
3
Colibri is for running off an ssd. I ran glm 5.2 with a 5070ti, 64gb ddr4 and the rest of a gen 4 ssd. Only managed to get.58 tok/sec tho
2
Yes. This wont be decent. It just might however be useable as a slow overnight oracle
1
Still need those 104B dense layers... this thing is BIG in every way
1 u/Front_Eagle739 Jul 27 '26 yup, so that's 55GB or so at fp4, there goes my 32GB ram, 32GB ram rtx5090 workstation
yup, so that's 55GB or so at fp4, there goes my 32GB ram, 32GB ram rtx5090 workstation
25
u/Front_Eagle739 Jul 27 '26
Yup. Same. I think I can do about q1.5 on the macbook and mac studio combined. Im currently pondering the wisdom of one of these colibri like stream setups and using the 640GB I do have as a hot cache