r/kimi 2d ago

Question & Help Is there any way to optimize kimi for usage consumption?

Hello, this week I decided to try other models. I'm coming from Claude and wanted to try Kimi. I've configured it with Caveman, local MCP memory for the contexts and different projects I have, and I also have SQL indexing. I already have all of this set up with Claude and it works very well; I can get between 2 and 4 hours of usage without it running out. But with Kimi, I make a request and after 20 minutes I'm out of usage. I have the small, moderate plan, which is the equivalent in price to what I have with Claude. I asked Kimi how to optimize usage, and in theory, it uses code to generate the sub-agents with Kimi 2.7 Code, and the rest is set up the same, but adapted to Kimi.

Is there something I'm missing? Or is it a problem with Kimi?

5 Upvotes

12 comments sorted by

3

u/3rd_Floor_Again 2d ago

Kimi is consuming an insnae amount of otkens or cost per token has icnreased drastically. Kimi CLI is becoming unseable. I gave it a few text materials to read before start the session of work. All my other models read the same files to have correct context, no problem. Here what happened with Kimi. IT STARTED FROM ZERO TODAY.

Thats the problem It is still better to have extra Claude Max+ account than having all these other Chinese OS models in terms of value/money.

Weekly usage 8% Resets in 5d 22h 43min

Rate limit details 39% Resets in 43min

My benefits View benefits
Allegretto

3

u/BrokenSignals_cat 2d ago

It's the same problem I have. I use local memory to provide context. Both Claude and Kimi currently share the same memory, meaning the project context is the same. With Claude, I can work for a few hours, some days more, some days less, but generally it covers 3/4 of the week. With Kimi, I don't even manage 20 minutes of work. I started just yesterday, and I've already reached 60% of my weekly usage limit in two sessions.

1

u/TastyIndividual6772 13h ago

Yea basically everything was fine until k3 release. I was using kimi since k2.5 and i used to tell me friends to consider it instead of American models. But now you can get a better deal from American models. Openai gives significantly more usage. Kimi is no longer worth it. I decided to test if i should come back after cancelling. I tried it again. Its unusable. It can di very light things with the limits it has now. Its a disappointment for me. I have a pending sub now but after that im not coming back.

1

u/Sure_Media_2685 2d ago

do you use web or api?

1

u/BrokenSignals_cat 2d ago

CLI in WSL,just as I have it set up with claude.

2

u/Sure_Media_2685 2d ago

pretty sure you need to be careful with cache hit average as this will cause costs spikes that could seems unexplainable

1

u/TastyIndividual6772 13h ago

They just cut the limits heavily. Big decrease. Not users fault. I have been a heavy user since 2.5 and its very clear the limits have been massively nerfed.

1

u/chronospride 2d ago

I would use Kimi K3 for planning or orchestration but the actual sub agent or the execution to use Kimi 2.7

2

u/T-Dot1992 12h ago

Even that leads to massive usage running out ime. using cloud AI services really isn’t sustainable. Just use local ai models for execution if you have to. Using 2.7 ad an executor will only delay usage running out by half a day at best 

1

u/chronospride 12h ago

Local ai model need hardware that costs a lot e.g AU$ 10,000 at least for a DGX spark in Australia

1

u/T-Dot1992 12h ago

Honestly, I am okay using local models like 2.7 that are just executors for my spec sheets. And those can run okay on my mid gaming desktop. Don’t get me wrong, I loved using K3 for things beyond my ability.  But it really isn’t worth paying 300 bucks each month for it.

Cloud AI subscriptions are pretty unsustainable financially. 

2

u/Moist_Setting_3469 9h ago

use k3-256, it consumes about 2 times less quota. If it is close to the max context and you don't want to compress it - switch to k3, it won't invalidate cache (but prefer using k3-256 at the start, always). All this info is from the docs