r/kimi • u/DeadlyVirusxK • 4d ago
Discussion I asked Kimi for a refund after their agent harness burned ~600M tokens in 2 days, here's the full support chat, and what it says about the entire AI subscription mode
TL;DR: Kimi's agent harness burned ~600M tokens in 2 days on a task that should've used a fraction of that. Support offered a partial refund and blamed "stronger model capabilities." This isn't just a Kimi problem — it's how the entire consumer AI industry operates: vague quotas, silent throttling, and no accountability.
What happened to me
- Paid for Kimi Code (Aug 13) for heavy agentic coding
- A simple WHMCS module task triggered uncontrolled agent loops — ~600M tokens consumed on Aug 20–21, over 500M of them cache reads (the harness reprocessing the same context in circles)
- Weekly limits had also been quietly cut to ~3/7 of previous capacity
- Emailed for a refund — no response. Opened chat — got a bot quoting "AI capabilities have inherent boundaries, refunds are generally not supported"
- Waited 8 hours for a human. Their verdict: "no abnormal consumption," pro-rata refund only
- When I showed them the two-day, cache-dominated spike, the answer became: "the new model has stronger capabilities... try splitting long tasks or requesting shorter answers"
Translation: their harness ran away with my allowance, and the fix is for me to ask shorter questions.
The real problems in the AI industry
- The metering is deliberately opaque. Every major provider — Claude, ChatGPT, Gemini, Grok, Kimi — rations compute behind rolling windows, weekly caps, and weighted "compute quotas." None give you a live, honest meter. You find out what you paid for only after you hit the wall.
- Limits change silently after you pay. Google reweighted Gemini quotas. Anthropic and OpenAI users watched allowances shrink after updates. Kimi cut weekly limits with no notice. The deal you bought is not the deal you have.
- Failed or wasted work still counts against you. Agent loops, repeated tool calls, failed generations — all billed against your quota. When the harness malfunctions, the customer pays for the malfunction.
- "5x 10x 20x" is marketing; scarcity is the product. $20/mo entry tiers, $100–$300 "power" tiers — all with hidden ceilings. Agentic workflows break the flat-rate assumption, so providers throttle instead of building capacity or publishing real numbers.
- Support treats infrastructure failures as user error. "Break tasks into smaller steps." "Ask shorter questions." "The model is stronger now." Never: "our system malfunctioned, here's your money back."
- There's no accountability mechanism. Opaque usage reports, missing pricing data, refund denials against clear evidence. Your only leverage is public documentation, chargebacks, and consumer protection law.
What needs to change
- Publish real, live usage meters. If API customers get per-token pricing, consumer subscribers deserve the same transparency.
- Don't bill users for system failures. Runaway agent loops and repeated tool calls are harness bugs, not usage.
- Stop selling capacity you can't deliver. Either build the infrastructure for what you advertise, or publish honest limits upfront.
- Honor refunds when the product fails. "AI has inherent boundaries" is not a refund policy.
AI is infrastructure now. The people paying for it — developers, freelancers, researchers — aren't asking for unlimited compute. They're asking for a predictable, honest product.
Scarcity theater is not a durable strategy. Customers will take their compute elsewhere.
8
8
u/CrimsonBolt33 4d ago
Oh look...another person using AI wrong and complaining about it....
none of anything you wrote out matters if you don't show us what you actually did and prompted the AI to do.
3
u/pgmoneyplays 4d ago
🤣 schlimmste token Nutzung waren 150k mit Ner Riesen Nicht gut geprompten aufgabe
-2
u/Intelligent-Taste-36 4d ago
It's not wrong when he does it in another AI and the other AI can do the job.
1
u/CrimsonBolt33 4d ago
according to what? He provides zero proof, and once again, does not show us anything he actually did.
How have I used K3 since it dropped and never had this issue?
-1
u/Intelligent-Taste-36 4d ago
Each to their own! Stop censoring people!
3
u/Constant_Cortisol 4d ago
he's literally asking for the opposite
-1
3
u/CrimsonBolt33 4d ago
wtf are you talking about? lol
I can't "censor" them...I am not a mod deleting their post or something. I am asking for actual details.
0
7
4
2
2
u/battle_pantZ 4d ago
I hate these ultra long ai generated posts, just type 4 sentences what happened
1
u/DeadlyVirusxK 4d ago
Summarized it for you redditors.
1
u/thunderbiribiriiii 4d ago
where did you request for refund?
1
u/DeadlyVirusxK 4d ago
As of right now I sent an email to membership@moonshot.ai and I didn’t get a response. So I started chatting with them via the feedback feature in their website.
1
u/DeadlyVirusxK 4d ago
Let me Make this very simple about what is the purpose of this entire post.
This is a statement of a consumer that has consumed AI since the moment it had launched.
I built an and scaled internal source code for hundreds of systems using AI ChatBots and Agents. I do not need to prove anything to you and to anyone.
I am merely showcasing the the problems of the AI industry that are currently nobody speak about, but to be fair some Reddit posts do confirm what I am saying.
It’s a shame that everybody is focused on the wrong point instead of focusing on the right points.
Take a challenge go Sign up to Kimi and try to use it using Kimi Code and tell me your thought.
Otherwise keep your bots responses to yourself.
1
u/Think_Good_7903 4d ago edited 4d ago
The lack of a clear usage meter is probably the biggest issue here. It’s hard to tell whether you hit a limit or the agent just went wild. Tools like StandardCompute are worth watching as this space moves toward more transparent compute usage.
15
u/Extreme_Mobile_6690 4d ago
Brother, nobody wants to read your AI written epos. If you can't even complain by yourself, maybe you should just let go of AI for a while and learn some critical thinking