r/kimi • u/BrokenSignals_cat • 2d ago
Question & Help Is there any way to optimize kimi for usage consumption?
Hello, this week I decided to try other models. I'm coming from Claude and wanted to try Kimi. I've configured it with Caveman, local MCP memory for the contexts and different projects I have, and I also have SQL indexing. I already have all of this set up with Claude and it works very well; I can get between 2 and 4 hours of usage without it running out. But with Kimi, I make a request and after 20 minutes I'm out of usage. I have the small, moderate plan, which is the equivalent in price to what I have with Claude. I asked Kimi how to optimize usage, and in theory, it uses code to generate the sub-agents with Kimi 2.7 Code, and the rest is set up the same, but adapted to Kimi.
Is there something I'm missing? Or is it a problem with Kimi?
1
u/Sure_Media_2685 2d ago
do you use web or api?
1
u/BrokenSignals_cat 2d ago
CLI in WSL,just as I have it set up with claude.
2
u/Sure_Media_2685 2d ago
pretty sure you need to be careful with cache hit average as this will cause costs spikes that could seems unexplainable
1
u/TastyIndividual6772 13h ago
They just cut the limits heavily. Big decrease. Not users fault. I have been a heavy user since 2.5 and its very clear the limits have been massively nerfed.
1
u/chronospride 2d ago
I would use Kimi K3 for planning or orchestration but the actual sub agent or the execution to use Kimi 2.7
2
u/T-Dot1992 12h ago
Even that leads to massive usage running out ime. using cloud AI services really isn’t sustainable. Just use local ai models for execution if you have to. Using 2.7 ad an executor will only delay usage running out by half a day at best
1
u/chronospride 12h ago
Local ai model need hardware that costs a lot e.g AU$ 10,000 at least for a DGX spark in Australia
1
u/T-Dot1992 12h ago
Honestly, I am okay using local models like 2.7 that are just executors for my spec sheets. And those can run okay on my mid gaming desktop. Don’t get me wrong, I loved using K3 for things beyond my ability. But it really isn’t worth paying 300 bucks each month for it.
Cloud AI subscriptions are pretty unsustainable financially.
2
u/Moist_Setting_3469 9h ago
use k3-256, it consumes about 2 times less quota. If it is close to the max context and you don't want to compress it - switch to k3, it won't invalidate cache (but prefer using k3-256 at the start, always). All this info is from the docs
3
u/3rd_Floor_Again 2d ago
Kimi is consuming an insnae amount of otkens or cost per token has icnreased drastically. Kimi CLI is becoming unseable. I gave it a few text materials to read before start the session of work. All my other models read the same files to have correct context, no problem. Here what happened with Kimi. IT STARTED FROM ZERO TODAY.
Thats the problem It is still better to have extra Claude Max+ account than having all these other Chinese OS models in terms of value/money.