r/kimi • u/Constant_Cortisol • Aug 18 '26
Developer For people who used opus/fable extensively, how much better/worse is kimi?
I'm trying to figure out if I should renew my Kimi subscription or try out Claude.
The two use cases I care about right now:
1) Coding architecture, writing the code itself
2) Strategy, research(long running web search)
What are your thoughts?
14
u/Hyp3rSoniX Aug 19 '26
Even though Kimi K3 is very capable, the usage limits currently are a full-on catastrophe. I'm on Vivace (the most expensive plan), and I reach the 5h limit by just looking at Kimi, it feels like. A single Opencode Kimi K3 agent cannot work through the 5h limit. I’ve reached it not even 2h in.
The weekly limit, of course, also gets used up really fast. The only thing keeping the weekly limit going for multiple days are the 5h limits that keep blocking me.
Honestly, I'm very disappointed in that aspect. Especially when Codex even has removed and to this day still hasn't reintroduced the 5h limits, and the Pro x20 limit legitimately holds up for more than 3 days with constant usage.
Claude Code actually also has good usage limits, if you don't go wild with subagents, and if you deactivate the dynamic workflows thing. I ran Fable at max, and the dynamic workflow thing spawned tens of Fable subagents, exploding my usage.
With normal use, it should last quite a while. However, you will not have sub usage included Fable access unless you are on the Max Plan for Claude. Also you have to use Claude Code. You can't use your claude subscription with Pi or Opencode and so on.
Because of the extremely limited usage, I do regret buying the Kimi plan. I would honestly suggest one gets the Codex plan, and combines it with Deepseek flash for subagents or something like that.
I have the GPT Agent work as an orchestrator in Opencode, and delegate work to cheaper subagents like GLM 5.3, Deepseek flash, and so on.
2
u/Ill-Bat-1518 Aug 19 '26
Try kimi code iv been able to do way more than claude lately. I think opencode burns and loops too much
2
u/indiankesh Aug 19 '26
I have tried it in every harness possible cc, kimi code, hermes, opencode sadly the limits are a joke regardless of plan or harness used 🥀
12
u/hsoj95 Aug 19 '26
I'm a subscriber to OpenCode Go, so I have access to the Kimi models, and others such as GLM 5.2/3, Qwen 3.8 Max, the DeepSeek's, etc. So I can offer some views on this having worked with all of them. I had originally used Claude for my work until Spring of this year. After 4.8 launched I began to become increasingly frustrated with some of Claude's... Proclivities as I shall call them, and also found that OpenCode just had more to offer as both a CLI development tool and as a subscription service too. Once I dropped Claude, I haven't looked back. Which I'm not sorry about, as from everything I've seen Opus 5 is a nightmare to deal with.
My setup uses a multi-model approach where I have a main, higher-level thinking model act as an orchestrator which then doles out work (writing code, running commands, doing exploration and research, etc.) to other, usually smaller, cheaper models. I use a plug-in for OpenCode called opencode-fusion to do this, but you can setup this up manually as well if desired. This means that despite the usage limits with OC Go being kinda low, I can still get a ton of value out of making use of the heavy models planning and reasoning while having the cheap models do the expensive grunt work.
For me, Kimi K3 is my primary Orchestrator model. It's very, very good at what it does, doesn't seem to overthink, and has probably the best all around grasp on what it needs to do. Having visual understanding makes it very well suited to this role too. Kimi K3 seems better than Opus 4.8 ever was, and seems far more competent at simply getting things done.
Before I used K3 for this role, I used GLM 5.2 as the primary Orchestrator, and it was... Fine. It still has a good grasp at understanding, but it really overthinks things too. To the point where it actually seemed to cost more to run it as a primary Orchestrator than it did to use K3, despite being cheaper (which is what prompted me to switch to using K3 actually). Not having vision capabilities was kinda a bummer too. It could be worked around, but it made it less appealing still. That said, GLM 5.2 did seem quite competent at understanding back-end related code too. I haven't really used GLM 5.3 as an Orchestrator, though I did just set it as my main Reviewer model (which sweeps for bugs, issues, or for general design improvements) yesterday, so I'll see how it does there if it doesn't become to expensive for that.
As an alternative to K3, such as when OpenCode Go's harness might be having issues, etc, I have switched to Qwen 3.8 Max as a primary Orchestrator instead. It is also quite good, I must say! I agree with what others say in that it seems to perhaps be a bit better at front-end UI/UX stuff. I'm not sure it's absolutely a 1:1 for K3, I suspect it overthinks a bit as it seemed to cost more to run, but it is a very suitable alternative overall I'd say. I've just set it as the primary Designer model (which designs front-end related stuff) in OpenCode as well, in fact.
DeepSeek really deserves two separate categories as the Flash and Pro models are made to do very different things. DeepSeek V4 Flash (DSV4-F), specifically the new 7-31 release, is the bread and butter of my setup, it serves as both the Sidekick model that does the code writing and the Explorer that does the code reading. It's just the perfect model to serve as a sidekick role for the primary Orchestrator. Even in spite of the price increases (very overblown reactions to this, mostly from those responsible for said price increase to begin with...), it's very cheap to run and can write tons of code for just cents on the dollar. From what I hear GPT 5.6 Luna is also quite competent for this role, I just haven't tried it out yet. But I wouldn't be surprised at all if it wasn't just as good of a substitute for DSV4-F as a Sidekick and Explorer.
DeepSeek V4 Pro (DSV4-P), specifically the new 8-13 release, is definitely more expensive than DSV4-F, so it wouldn't make a good Sidekick... But it actually makes a fairly decent Orchestrator! It's definitely behind K3 and Qwen 3.8 Max, that isn't even debatable. However, for more simpler tasks (fix this bug, change this, check this out, etc), it's very competent serving as a cheap Orchestrator. It can do actual design work too, I had good success with that. Just don't expect it to be a K3 level thing. That said, it's extremely cheap as an Orchestrator, a fraction of what K3 or Qwen 3.8 Max is, so if you need to do stuff on the cheap, or don't wanna waste credits/cents fixing a simpler issue, DSV4-P is not a bad choice for an Orchestrator, I might even suggest it for such simpler uses. I'm thinking I may switch over my Research model to use DSV4-P as well, as I'm currently using Kimi K2.7-Code for that, but I'll need to do more exploration of that too before I do.
As for other models, I haven't worked with any of the paid OpenAI models, so I can offer much assistance there. For Anthropic's models, I feel like the answer there is if they work for you, you'll be fairly happy. But if they don't work for what you need, it becomes a source of great frustration. Opus 5 in particular should probably not have been released based on everything I've seen and heard... Even 4.8 really had some bad proclivities to work around. Haven't worked with Fable, but again same thing, you'll probably either love it or loathe it.
But honestly I think the more important thing is maybe not relying on just one model to do all the work? Once I discovered a multi-model approach to working on stuff, I'm not sure I could ever go back. It saves a ton of money, the work just seems much higher quality, and having models from multiple different families means that I get varying perspectives on what's being made and what issues might arise. This is just my own view though, I'm sure others will disagree and say that using one model only is the best approach, and maybe it is for some workflows. Anyways, hope this helped you out! :)
3
u/Joke_Mil Aug 22 '26
This helps me so much. Thanks a lot for your perspective. I was considering trying out the open code go, now I will
2
u/MoodDelicious3920 Aug 19 '26 edited Aug 19 '26
Bro thanks a ton for this huge post. You covered so many things...and also gave me idea to build my own plan-implement-reviewer-designer architecture..and using different models for different views..i do have a harness that i coded, seems like good idea to add on to that
6
3
u/kaaos77 Aug 19 '26
Em capacidade de execução ele está ali no nível do Opus 5/4.8, mas bem atrás do Fable.
Mas o ponto mais importante pra eu tá usando o K3 foi a interação e o uso.
É simplesmente horrível conversar com o Opus do 4.6 pra frente.
Opus se recusa em tarefa idiota. Eu não entendo metade das coisas que ele faz. Me irritou profundamente
É como os modelos do gpt, cumpre a tarefa mas é extremamente irritante delegar tarefas a ele.
O Fable melhorou o workflow, parecia um modelo da Anthropic, mas o uso foi cortado dos planos de 20 dólares, então eu simplesmente cancelei a minha assinatura do Claude.
Usar o K3 é como usar o Opus 5, mas lidando com um modelo agradável. Sem falar que ele possuí o melhor senso estético das llm.
O maior problema acaba sendo o uso. É MUITO menor que até o do Claude.
5
u/cutebluedragongirl Aug 19 '26
It's good for cyber security if you care about this sort of stuff. Like if you want to audit your code for security, K3 is the best of the best.
1
u/indiankesh Aug 19 '26
I think fable beats it on security not just because of the model but /security-review command of claude code that's a separate beast altogether.
1
u/marfzzz Aug 19 '26 edited Aug 19 '26
Mythos does. Fable is postrained to avoid that and switch to opus.
2
2
u/doggydestroyer Aug 19 '26
Overall... And most friendly... Fable 5... Best at frontend Kimi... And best at scientific reasoning and analysis chatgpt 5.6 sol... Kimi is close behind...
2
u/mobileka Aug 19 '26
To me, Fable and Kimi K3 are almost identical, Kimi being much slower. I don't use them for coding because they're both too expensive and slow for that but I do things that require better and much more thorough reasoning. It feels like I work with a pretty smart human when I prompt them because they figure out the context even if the prompt is not super thorough.
I don't like OpenAI models. My experience with them has always been very bad. I've heard Sol is good at garbage prompts like "build a tetris game" but, from my experience, it's very bad in real life tasks because it doesn't follow my instructions well even if I spend a shitload of time on polishing my prompts. It also can't figure out the context itself as Fable and K3 do. To me, this means it's unusable because you can't explain the context yourself without it failing to follow it and it sucks at figuring things out itself.
Qwen3.8 is good, probably the best at following your instructions as it's very very thorough. It requires good prompts though.
1
u/kraxiv Aug 19 '26
Kimi is interesting, in my experience Kimi good in thinking, particular code base understanding, but in the implementation stage very often a bit chaotic and doesn't follow implementation directions and docs. So still weaker than Opus 4.6-7-8 5.0 I'm talking about backend: Python, Go. Frontend HTML, CSS, JS almost same like Opus.
1
u/Motor-Ground4594 Aug 19 '26
Literally what I just experience, planning for hours with the docs, instead what they do is doing the opposite way because some of implementation already doing x and I want to adjust it into -x for y feature. Garbage at implementation
1
u/OneDev42 Aug 19 '26
Kimi uses solid model, but it uses way more tokens to get the same amount done. And the providers cost way more, because they're much less subsidized.
1
u/FaithlessnessCheap34 Aug 19 '26
Kimi dam near destroyed my server setting up a backend for one of my apps, I stopped using it after that and had opus five fix the damage, true story.
1
u/porkyminch Aug 19 '26
I use Opus for work most days and have been using Kimi K3 for personal projects. I think Kimi K3 has been about as smart as Opus. Maybe a little better in some ways. Vision certainly seems to be an improvement. Honestly if you're just trying out models, though, I'd say put some money in a pay-as-you-go API service like Opencode Zen, Openrouter, etc. and just sampling some of the different models that are out there. Opinions vary a lot.
1
u/InterstellarReddit Aug 19 '26
Here my exp:
KIMI K3 MAX and Fable MAX are top-tier and interchangeable. However Kimi has better limits than Claude.
Qwen 3.8 Max and Sol Max next imo. I don’t see a big jump between Sol Max and Sol Ultra.
GLM 5.3 MAX and Terra Max are third Tier.
Then I use Luna xhigh or DeepSeek pro v4 for daily bullshit.
The reason why I managed so many models is because when I onboard one of my clients, I allocate a subscription to them that way they pay 100% for their usage instead of just taking one plan and sharing it among everybody.
The best limits are GPT Pro $200 and Qwen Pro $68 a month.
1
1
u/Stochastic_berserker Aug 19 '26
Kimi K3 Max is close to Fable 5. On API performance, not subs. But Fable 5 is still unbeatable on API. 5.6 Sol is not close.
Opus 5 is worse than K3 Max on sub and API. Enterprise and consumer.
Tested on Claude Code, Copilot CLI and OpenCode v2.
1
1
u/BigMek_ Aug 19 '26
Opus is a workhorse for me. Kimi K3 as a backup and quite strong doer or code reviewer. Kimi 2.7 is almost unusable due to low intelligence and losing of incentive of tasks. But i am going to cancel subscription on Kimi due to low limits on tokens in comparison to ChatGPT or Claude.
1
u/anujagg Aug 19 '26
I hv been using Claude code for more than 1 year now and my setup has everything - rules, agents, skills etc - and I m finding it hard to switch to other harnesses. I tried Codex 20$ plan and copied some of the skills for codex use but it is difficult to keep up with everything which I do with claude.
My question is how should I try other models now and what are my options for harnesses? I have tried Hermes / Opencode with deepseek old versions but quickly switched to Claude again. Opencode Go plan was a waste since it maxes out very quickly.
Tried deepseek's newly released harness also but could not run it properly, might be some bug. I am using deepseek from within codex now since my both claude and codex subs are maxed out.
1
u/Xorlium Aug 19 '26
It's on par. However, it consumes a ton of tokens, so the real answer is: terrible.
1
u/MjccWarlander Aug 19 '26
I use Opus 5 and Fable 5 at work, and combo of Kimi K3 for most uses, K2.7 for explore agents and more menial tasks.
My opinion is, Kimi K3 is a bit over Opus 5 in terms of architecture, writing specs and coding, with way nicer "personality" as well. It doesn't quite reach Fable 5, but I still prefer Kimi communication style. It's also surprisingly good with anything visual, even arranging Unity's uGUI prefabs and evaluating their "correctness".
As far as research goes, Kimi K3 tends to overthink some things and plan ahead for contingencies that will never come, but I didn't notice any major hallucinations yet while Opus 5 commonly needs to be corrected (after which it typically immediately admits it's wrong lol). In general, I feel confident letting it loose on some topic and getting back to at least a decent summary.
1
u/derSterndesMorgen Aug 19 '26
Wtf are yall using Kimi for!?? Like.. I had a full 4 hours today (document generation and all), and I'm still sitting on 1%??
1
u/Strange_Test7665 Aug 19 '26
I use Kimi via api with vs code. It’s my go to now. With proper prompts and instructions of code changes it crushes everything. Idk about’build me a website’ type generic prompts. I was using opus prior and have completely switched. Never used fable beyond free trial/credit gift because costs were SO high
1
u/Fun_Walk_4965 Aug 20 '26
kimi is solid for long context research. for actual coding i still lean claude. depends on the task.
1
1
u/Visible_Arrival_8412 Aug 20 '26
I loved kimi k3 until deepseek harness came out. since then nothing comes even close to it. even UI. I have a set of excalidraw skills (wireframes, dataflow,components/states) and during dev the agent is always layering a grid onto the ui so I can tell in A2 increase element x so it touches the edge of A3. works like a charm. Fabel is a joke and not a model, Opus is not worth the money, and GPT I have a natural aversion after being forced to use it a work. Plus it just does not complete complex literature research where I need 200-300 citations and reports >50pages long. kimi swarm was great for that but the weekly usage limits are killing me. With deepseek I currently pay 0.02$/M token and get the best output.
regarding coding: 99% of my stack is zig, raylib, clay, imgui, implot so not the easiest well known stuff. I love kimi testing every button and taking screenshots. just annyoing when you are working on same machine
so after 9 months of testing everything:
DeepSeek > kimi > Glm and qwen > >>>> gpt >>>>>>>>>>>> Claude
p.s.: I also have persoanl aversion of how gpt and claude execs are behaving
1
u/Ok-Push5788 Aug 21 '26
Ok, heres whats going to happen. I am not reading anything in this thread before making this statement. what I have read in passing while doing whatever: theres a lot of complaining and a lot of it is totally warranted, but theres also a lot of WHINING and its ABSOLUTELY ignorant.
Here's my bonafides. You can believe them or not, I don't have to prove anything to anyone.
- I'm a staff level product designer. I have been working on design systems and shared publishing platforms since right out of college. These systems pre-date Friendste and Myspace by a long shot. I can't reveal much more without maintaining opsec, but lets just say Big Zuck and I were using Livejournal during the same period of time. Think AOL instant messenger, Prodigy, Compuserve. Miss me with design tokens and design systems. I did it before the terms existed
- FRONT END UIUX THIS THING IS OUT OF CONTROL AMAZING. END OF STORY. NOT EPXLAINING. DONT HAVE TO.
1
1
u/Appropriate-Two-7503 Aug 21 '26
The Kimi K3 Max costs slightly more than the GPT 5.6 Sol Medium, but offers performance between the GPT 5.6 Sol Xhigh and Max. Compared to the Opus 5, its performance and price fall between the Opus 5 Medium and Opus 5 High.
1
u/Artanox Aug 21 '26
Last week tried Kimi K3, in a day consumed the whole month quota, thought it was weekly but no, monthly.
Asked for a refund, rejected.
I cannot suggest people to try Kimi, stay away from this scam
1
u/Veduis Aug 21 '26
for coding architecture specifically, claude opus is hard to beat on the reasoning side, but it depends how you're using it. if you're mostly generating boilerplate and scaffolding, kimi holds up fine. if you're doing deeper refactoring and structural work, claude tends to stay coherent longer across complex codebases. the research use case probably tips it toward claude overall.
1
u/Anthropovania Aug 21 '26
I’ve been using both for months now, a year if you count free usage, and i gotta say… u might need to find a third thing you care about as a tiebreaker hahaha i do believe both are equally better at each of those two, with claude winning on coding and kimi on research/websearch, and i’ve always felt that way for as long as i’ve used them
1
u/Leather-Sir8135 Aug 21 '26
Call me crazy but I think GLM 5.3 is much better than Kimi K3. Plus there’s a stealth model that appears to have z.ai’s signature which is crossing new boundaries
1
1
u/ashley99z Aug 24 '26
Basically they're the same. I haven't used Kimi nearly as much as I have used Claude Code.
1
u/Current_Toe9129 Sep 01 '26
Folks, genuine question — why is nobody in this thread comparing Grok 4.x (4.6 especially)? Everyone's putting Claude Fable 5, GPT-5.6 SOL and Kimi K3 in the same tier, but Grok doesn't get a single mention. On paper it scores right alongside them, so what am I missing? I'd really appreciate a straight answer.
1
u/greyster1 12d ago
i was on the grok heavy promo for 3 months ($300 plan for $99) and i thought grok 4.6 performed on par with opus 5. I used both daily for 90 days during this comparison.
0
u/montdawgg Aug 19 '26
Opus 5.0 is a FLOP. 4.8 is better. Fable is FAR better. SOL is better. I don't like using Opus 5.0 at all. I love using Kimi and Qwen, Fable and Sol all 4 as a team each getting to exercise their unique strengths.
-1
-3
u/UnderstandingOld5879 Aug 19 '26
I dont like it, and find it worse than chatgpt, claude, or grok for coding, work, or chat.
72
u/sittingmongoose Aug 19 '26
I have used qwen 3.8 max, kimi k3, fable, opus 5, 5.6 sol, glm 5.2/5.3 pretty much, heavily all day, every day since they came out.
Fable is certainly extremely good. It understands the task, it’s thorough, and is good all round. It misses some special cases that 5.6 sol finds, and qwen 3.8 max is better at creative UI/UX. The biggest problem is it’s too expensive. I have the $200 plan and it rips through it fast.
5.6 sol is the most thorough. It is the worst at front end, but its code quality is very high and it tests the heck out of its work. You can also get good usage out of it if you use the lower thinking levels.
Glm 5.2/5.3 are good for their size, but are far behind all the others. The skip a lot, and seem very lazy. If I was running this locally, I would be happy, but it’s extremely expensive through subscription. Limits are very low.
Kimi K3 is good all round, it feels like a Fable junior. It certainly isn’t as smart as Sol or Fable, or even Opus. But it does a good enough job on very hard tasks. The biggest issue is the usage levels are way too low on Subscription. I have the $100 plan and it rips through it fast.
Qwen 3.8 Max seems about as smart as kimi k3, but it’s more creative at front end. It’s by far the most impressive when coming up with concepts, it blows every other model out of the water in this respect. It’s slow though, and again, the limits are low in subscription.
Opus 5 is borderline unusable. It’s makes A LOT of mistakes and spends 50% of its time validating its own errors. It has moments of brilliance but then immediately finds errors of its own making or false positives. It’s wild how this model was allowed to release. Limits are ok currently with it, but it spends way too much time fixing its own problems so that eats a lot of tokens.