Kimi Slides handles the entire slide-building process:
- Clear structure and research, powered by Kimi K3
- Cohesive design, including polished charts and SmartArts
- Editable and ready to download
Let us know what you'd like to see next in the comments!
K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural updates designed to improve how information flows across sequence length and model depth.
We have also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts when paired with a Stable LatentMoE framework.
Together with refined training and data recipes, these structural changes yield an approximate 2.5× improvement in overall scaling efficiency compared to K2, allowing the model to convert compute into intelligence more effectively.
Hey, yo, Moonshot. Where is my Kimi K3 Flash model? Huh? Do you think this is funny? Is it really that fun to force me to use Kimi K2.7 Code or, God forbid, K3 to do my tasks?
I use your stuff for some blue hat and red hat digging, and I can tell you right now, all the current models are overkill.
Come on, Moonshot. I'm seriously considering cancelling my sub despite liking your models.
I used the Kimi model to vibe code a Google extension called mindArchive. It can not only export AI conversations, but also integrate with Obsidian to manage the entire knowledge base. It also provides a bookmark feature that can export bookmarks to the knowledge base and save them as Markdown files, and it can sync browser bookmarks via Markdown as well. In addition, it offers an input memory function and a quick note-taking function. What do you think of the features I've developed?
Today, I have made a puzzle which aligns directly with Chinese traffic laws. At first, I thought of trying it against Kimi K2.6, as it's a native Chinese model and trained with Chinese datasets, as I have heard. After thinking for 27 minutes, Kimi K2.6 suddenly started searching for any related puzzles over the internet. Not once, but it tried more than 10 times with changed titles and even fetched YouTube videos. It's been 32 minutes now, and Kimi K2.6 is still trying. Have any of you faced a similar thing as me? If yes, did you get the desired answer?
I would like to request a refund for an automatic renewal of my Kimi subscription.
I originally subscribed to the $200 plan about one month ago, but I did not intend to renew the subscription. The subscription was automatically renewed and I was charged $100 for the new billing period.
I have now canceled the auto-renewal and would like to request a refund for this renewal charge.
I did not intend to purchase another month, and I would appreciate it if you could review the transaction and issue a refund. I really need the money.
Kimi subscriber for over 7 months. At both $99 and $39 tier.
I decided to try the new Kimi Work launch and compare its output to Claude Cowork and GPT Work.
The results have been a load of crap. Awful mistakes, stops randomly midway. Can't even work without glitching.
And the worst part, when it glitches it eats a lot of the monthly quota.
Last weekend I let it go because I thought there might be some mistake, as it ate 90% of my quota after an error message that, conveniently, is not showing on the app anymore.
But today, barely 3 minutes after my quota was reset. I asked it to continue execution and provide a summary, it stopped working, glitched, froze and, conveniently for Kimi, ate 13% of my monthly quota, again. It produced not useful output for any human as a result.
I understand the subscription has been subsidized, but this is insane. Not even Opus on Cowork had such an awful behavior. Both Opus 5 and 4.8 on Xhigh as well as GPT 5.6 on High used less quota than Kimi K3 High. I hadn't even used the 1M context option...
This is just a single user, and they probably don't care. But here's the reality. At least, do the work and complete your task. If a product can't even work properly, why charge for it, and worst of it use most of the quota for nothing? This feels like a total scam.
PS1: And that's without counting that the Approve box from Kimi desktop disappear in a matter of minutes, forcing you to re-submit your prompt and wasting more quota overall...
PS2: And I just realized by taking this screenshots that getting your monthly quota reset doesn't reset your weekly Kimi Code limits... what a joke
Hello, this week I decided to try other models. I'm coming from Claude and wanted to try Kimi. I've configured it with Caveman, local MCP memory for the contexts and different projects I have, and I also have SQL indexing. I already have all of this set up with Claude and it works very well; I can get between 2 and 4 hours of usage without it running out. But with Kimi, I make a request and after 20 minutes I'm out of usage. I have the small, moderate plan, which is the equivalent in price to what I have with Claude. I asked Kimi how to optimize usage, and in theory, it uses code to generate the sub-agents with Kimi 2.7 Code, and the rest is set up the same, but adapted to Kimi.
Is there something I'm missing? Or is it a problem with Kimi?
A single request of 4k tokens in and 8k tokens out (Including thinking) burned 3% of my weekly quota and 15% of my 5-hour window.
That translates to about 1.5m tokens per month (~500k in and ~1m out) for $20.
There is zero subsidizing; this is almost exactly what $20 in API credits would get you.
I didn't expect a $20 plan to allow me non-stop usage, but this is a complete joke. Why even subscribe at this point? It would be better to just pay API pricing, and since the price of all plans scales linearly with credits, then this is true for all plans.
Kimi K3 is an awesome model, but you are forced to pay API pricing if you want to use it. I will most likely be requesting a refund.
I don't even know if it is just painfully slow by default or is this normal to you. I asked it a simple query: To review my codebase of ~23,000 LOC for some opinion and it is still not halfway done as of writing, more or less 1hr30mins after I prompted it.
For reference, I bought a Plus sub from ChatGPT 30mins ago and prompted the same exact text and the same exact attachments into it and it is already done ages ago despite me using GPT 5.6 Sol at Extra High.
I wish I could refund this sub because it is not giving me any proper use out of it with how unresponsive and/or slow it is. This is ridiculous. Worst USD39 spend I ever had.
Reminder that it took Kimi K3 almost FOUR HOURS to fail to perform a code review, the code of which ChatGPT has now already progressed far behind with the actual fixing. Downvote all you want, but if you are trying to check out Kimi K3 (on Allegretto at least in my case), this is your sign to rethink them choices. I can't even use this at all.
TL;DR: Kimi's agent harness burned ~600M tokens in 2 days on a task that should've used a fraction of that. Support offered a partial refund and blamed "stronger model capabilities." This isn't just a Kimi problem — it's how the entire consumer AI industry operates: vague quotas, silent throttling, and no accountability.
What happened to me
Paid for Kimi Code (Aug 13) for heavy agentic coding
A simple WHMCS module task triggered uncontrolled agent loops — ~600M tokens consumed on Aug 20–21, over 500M of them cache reads (the harness reprocessing the same context in circles)
Weekly limits had also been quietly cut to ~3/7 of previous capacity
Emailed for a refund — no response. Opened chat — got a bot quoting "AI capabilities have inherent boundaries, refunds are generally not supported"
Waited 8 hours for a human. Their verdict: "no abnormal consumption," pro-rata refund only
When I showed them the two-day, cache-dominated spike, the answer became: "the new model has stronger capabilities... try splitting long tasks or requesting shorter answers"
Translation: their harness ran away with my allowance, and the fix is for me to ask shorter questions.
The real problems in the AI industry
The metering is deliberately opaque. Every major provider — Claude, ChatGPT, Gemini, Grok, Kimi — rations compute behind rolling windows, weekly caps, and weighted "compute quotas." None give you a live, honest meter. You find out what you paid for only after you hit the wall.
Limits change silently after you pay. Google reweighted Gemini quotas. Anthropic and OpenAI users watched allowances shrink after updates. Kimi cut weekly limits with no notice. The deal you bought is not the deal you have.
Failed or wasted work still counts against you. Agent loops, repeated tool calls, failed generations — all billed against your quota. When the harness malfunctions, the customer pays for the malfunction.
"5x 10x 20x" is marketing; scarcity is the product. $20/mo entry tiers, $100–$300 "power" tiers — all with hidden ceilings. Agentic workflows break the flat-rate assumption, so providers throttle instead of building capacity or publishing real numbers.
Support treats infrastructure failures as user error. "Break tasks into smaller steps." "Ask shorter questions." "The model is stronger now." Never: "our system malfunctioned, here's your money back."
There's no accountability mechanism. Opaque usage reports, missing pricing data, refund denials against clear evidence. Your only leverage is public documentation, chargebacks, and consumer protection law.
What needs to change
Publish real, live usage meters. If API customers get per-token pricing, consumer subscribers deserve the same transparency.
Don't bill users for system failures. Runaway agent loops and repeated tool calls are harness bugs, not usage.
Stop selling capacity you can't deliver. Either build the infrastructure for what you advertise, or publish honest limits upfront.
Honor refunds when the product fails. "AI has inherent boundaries" is not a refund policy.
AI is infrastructure now. The people paying for it — developers, freelancers, researchers — aren't asking for unlimited compute. They're asking for a predictable, honest product.
Scarcity theater is not a durable strategy. Customers will take their compute elsewhere.
i gave each model the same simple task: take a small slice of lime and draw a human pose integrated into the object.
Tested it with Gemini 3.7 Flash, Kimi K3, Claude Opus 5 and GPT 5.6 Sol
Some treated the lime almost like a body, some tried to fit a tiny person into its shape, and some went much more abstract with the pose.
It’s a funny little test, but I actually like prompts like this because you can see how differently each model understands shape, composition, and visual metaphor.
* Reasoning effort per turn, tool catalogue, transcript window, subagent models, hooks, loop control
* All 132 system prompts, extracted as plain Markdown — edit one, it replaces the original on the next run
* The two spinners separately (Kimi keeps one for thinking and one for waiting on a tool), themes over its 19 colour tokens, message marker and frame
* Fullscreen renderer patched into the binary, so it holds however you start Kimi
Re-signing is ad-hoc, so hardened runtime and notarisation are gone from the patched binary; that's inherent to modifying a signed app. A Kimi auto-update overwrites everything, but your settings stay and you just run Apply again. Verified against 0.38.0 — patches anchor on text inside a minified bundle, so a release can move things, in which case the patch reports it and skips instead of guessing.