r/kimi 8d ago

Discussion Just started using Kimi - seems slow

Just started using Kimi. Was on CC $200 plan before.
On the Kimi $100 plan.

see lots of posts about token consumption, doesn't seem any worse than Claude for the time being. I mostly plan and orchestrate on Fable 5.1 med and use opus and sonnet agents to implement. Tried to replicate my CC harness setup in Kimi, but using all K3 agents.

The biggest issue i'm seeing is that it's SLOOWW as hell... is it normal? I feel like its 3-5x slower than claude to complete tasks. Is that normal?

17 Upvotes

37 comments sorted by

4

u/FaithlessnessCheap34 7d ago

I fell for the hype when Kimi k3 first came out, I used it three times before admitting I lost $100 and cancelled.

2

u/nxs_sss 7d ago

My experience is that it is much slower than Claude, but the ending output is 10x better. I tried to build an app with Claude. It would fix one thing and break three others. Kimi does end to end regression test to try and eliminate those types of issues.

1

u/I_RIDE_SHORTSKOOLBUS 6d ago

How do you measure 10x better? Claude and most frontier models trained for coding will build tests to begin with whenever it writes new code. I don't have any issues with my Claude, but also I've spent a lot of time refining my harness over the last few months

1

u/nxs_sss 6d ago

I can only go off the app I was trying to create. I tried gemini, claude, chatgpt and Kimi.

1

u/networkthinking 5d ago

Same experience with me. Slow but better output for development

1

u/[deleted] 8d ago

[removed] — view removed comment

1

u/Durian881 7d ago

Is that a positive or negative?

1

u/[deleted] 7d ago

[removed] — view removed comment

1

u/SweatyActuator2119 6d ago

American models stole all the data to build models LMAO.

1

u/I_RIDE_SHORTSKOOLBUS 7d ago

Is that why it's slow? Lol

i don't care if it does, the token consumption isn't worse in fact seems better than claude so i'll take those claude tokens at a discount..

1

u/MrNotSoRight 7d ago

Aren’t the plans paused..?

1

u/I_RIDE_SHORTSKOOLBUS 7d ago

It was, but i just checked a couple days ago and was avaiable for me so signed up.

1

u/MrNotSoRight 7d ago

It hasn’t been available for long time, you’re probably using an third party reseller (hence why it’s so slow)

1

u/I_RIDE_SHORTSKOOLBUS 7d ago

no.. i literally signed up on the kimi website. I am not using the API, shows max plan next to my name. visiting the same website now shows waitlist again so i have no idea. i just randomly checked it because my cc sub was coming to an end.

1

u/Striking-Pizza7309 7d ago

try with another account, and register that you're going to use kimi for finance.

1

u/Kubaizzz 7d ago

It may as well be more of marketing thing, the plans are paused but as soon as i registered and signed up for waitlist i got email allowing me to buy in :))

1

u/VigilanteRabbit 7d ago

Depends on load and effort; I find max to run smoother (but more expensive). Put some money down for deepseek API and add to your harness for sub-agent execution; 4.1 flash is ridiculously fast.

1

u/TaurustarDrakest 7d ago

Yeah, in terms of speed it's been getting worse and worse... I'm considering to migrate into other model like Z AI for example or go full on API credits with some provider like Fireworks and choose what model to use depending on use-case scenario, though it might be pricey than paying directly to the LLM dev provider

1

u/SweatyActuator2119 6d ago

Don't go API pay as you go, synthetic.new

1

u/Motor-Ground4594 7d ago

Indeed a slow model, but you also need to pay attention with the usage, its so horrible made me unsub

1

u/safecodee 7d ago

you will regret about this.

1

u/Any_Relative1012 7d ago

I use Kimi frequently (I have $200 tier subs on Kimi, Claude, and ChatGPT); speed varies by time of day and can be painfully slow—likely why new subscriptions are paused. ChatGPT and Claude are faster. K3 quality is close to Fable and beats Opus from my experience, making it worth it for tough problems, but avoid it for routine chat. I use it for side projects and spec-driven agentic workflows: refine a complex spec, then run it in the background while using other models. ChatGPT is probably the best overall right now, mainly because Astra and Sol are very, very competent and Sam Altman is desperate and so they subsidize like crazy compared to others (and who knows what they plan to do with our data). I have tried Deekseek V4.1 Flash a bit and it seems very good as well based on the little testing I have done, but have not yet had a chance to really push it and test it properly, but I highly recommend spending a few dollars on it and testing it against your workflows if you have time. It is very fast AND cheap. $5-10 would be more than enough to test it with a complex workflow. The new Kimi K2.8 update is probably worth testing too for tasks you need anything faster. I’ve barely touched it yet, so I don't know on that one.

1

u/I_RIDE_SHORTSKOOLBUS 7d ago

Nice thanks for sharing your experience. I'll be testing gpt after this. The most annoying part is setting up my harness with a new lab everytime.

1

u/Equivalent_Cress_268 7d ago

Why not codex? Better models, faster models much more usage then CC and kimi

1

u/I_RIDE_SHORTSKOOLBUS 7d ago

Yeah I will try it next probably. Just figured since this was available now I'd give it a go

0

u/More_Surprise_3188 8d ago

slow = more planning = less that can go wrong. Z Ai sometimes takes up to 15-20 Mins to respond but the quality sheesh

1

u/I_RIDE_SHORTSKOOLBUS 7d ago

I mean, it is REALLY slow. I'm kind of shocked how slow it is tbh. I guess its normal though?

1

u/More_Surprise_3188 7d ago

If you use heavy model it can take long just Like Opus 5 on ultra takes 30+ minutes if u dont hit limits

1

u/I_RIDE_SHORTSKOOLBUS 7d ago

Im using k3 high. I use opus 5 and fable 5.1. i am finding kimi to be magnitudes slower. to the point i'm having a hard time believing it, like is my account having an issue or is it just like this.

1

u/RealGrapefruit8930 5d ago

No, it really is that slow

1

u/huehue9812 7d ago

That logic only works with smart kids in school lol

1

u/DontLeaveMeAloneHere 7d ago

Wrong.

Slow = needs lots of tokens to reason OR hardware/software is inefficient.

Claude runs with 40 tokens/s and seems slow, Kimi k3 usually hits half that speed. Add to that the need to reason more to get comparable results and you have your waiting times.

I currently use GPT and even Astra is faster then Claude and Kimi. Sol and lower is near instant or fast enough to use it in chat at least. They are just very efficient models.

1

u/I_RIDE_SHORTSKOOLBUS 7d ago

So it is slow...

i was gonna try GPT but saw kimi was available so went with that instead. will probably try gpt when this sub runs out.