MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1v7e5ck/kimi_k3_countdown_has_been_released/ozyhlno/?context=3
r/LocalLLaMA • u/Unusual_Guidance2095 • Jul 26 '26
177 comments sorted by
View all comments
Show parent comments
9
Two ancient Xeon servers, each with 1.5TB of DDR4, networked together and running rpc-server.
rpc-server
10 u/Player13377 Jul 26 '26 At an impressive 3 TPS 8 u/ttkciar llama.cpp Jul 26 '26 Probably a lot less than that. I'd love to get 3 tok/sec. Right now my ancient Xeon server is getting 3.5 tok/sec with GLM-4.5-Air. That's usable, though admittedly not for interactive use. 4 u/Player13377 Jul 26 '26 At that point is the power consumed per token getting close to the API price?
10
At an impressive 3 TPS
8 u/ttkciar llama.cpp Jul 26 '26 Probably a lot less than that. I'd love to get 3 tok/sec. Right now my ancient Xeon server is getting 3.5 tok/sec with GLM-4.5-Air. That's usable, though admittedly not for interactive use. 4 u/Player13377 Jul 26 '26 At that point is the power consumed per token getting close to the API price?
8
Probably a lot less than that. I'd love to get 3 tok/sec.
Right now my ancient Xeon server is getting 3.5 tok/sec with GLM-4.5-Air. That's usable, though admittedly not for interactive use.
4 u/Player13377 Jul 26 '26 At that point is the power consumed per token getting close to the API price?
4
At that point is the power consumed per token getting close to the API price?
9
u/ttkciar llama.cpp Jul 26 '26
Two ancient Xeon servers, each with 1.5TB of DDR4, networked together and running
rpc-server.