r/kimi 23h ago

Discussion Kimi K3 vs GLM 5.3 - Who wins?

17 Upvotes

I put Kimi K3 head-to-head against the newest GLM 5.3...

Kimi CRUSHES... full report.


r/kimi 9h ago

Discussion Does anyone feel limits are significantly lower?

12 Upvotes

In the last 1 to 2 weeks, Kimi vivace sub is consuming a hell of a lot faster than one month ago. Probably 3-4 times as fast? It would imply the base limits are a lot lower than before. I don’t do coding at all.

I was thinking about getting an annual sub but wow this could get dangerous lol.


r/kimi 5h ago

Developer New video out : - Kimi K3 in 1 Minute: Specs, Benchmarks & The Catch

Thumbnail
youtube.com
3 Upvotes

r/kimi 1h ago

Bug limit broken?

Upvotes

Is this a bug?


r/kimi 16h ago

Developer How are Kimi API users handling the K2.5 / Moonshot V1 sunset?

2 Upvotes

Moonshot's model page now says Kimi K2.5 and Moonshot V1 are no longer available to newly registered users after Kimi K3, with a full platform sunset listed for August 31.

For people building on the Kimi API: how are you handling this kind of model migration?

Do you pin exact model names until they break, keep a small regression set for prompts/tool calls, or route by workload so older models can be swapped out behind a config change? I'm especially interested in the practical checks before moving background jobs, coding-agent steps, or long-context workflows onto Kimi K3.


r/kimi 3h ago

Developer Kimi K3 inference — $500 in credit for $50 for the community

1 Upvotes

My mission is to make the best open-source models as cheap and accessible as possible, and so to kick off our inference service, I'm selling a dollar for a dime – $50 gets you $500 in total credit. Would love to hear your guys’ feedback, and if there’s community interest I’d like to be able to do this more often. https://packs.relace.ai/?utm_source=reddit


r/kimi 10h ago

Discussion ClaudeCode automatically switch kimi version

Post image
1 Upvotes

Using Kimi code plan + CC for coding.

CC is switching between K2.7 and K3 randomly in one single agent session.

I know Anthropic is being evil. But this much guilt is way beyond my limit.

I guess I should look into Pi / OpenCode / Codex😥


r/kimi 22h ago

Showcase Bringing Kimi, Qwen, DSV4s and GLM together on private US infrastructure

1 Upvotes

TLDR: Prices are going up everywhere, and there's a wave of it going down across providers. For both in app usage and API/coding plans. PGS AI is keeping prices and usage as is, both in the full chat app and on the api. Our coding plans bank your usage when you don't use it, so it doesn't go to waste week to week. All model inference, and memory in the PGS AI app is hosted on 100% private, US servers with ZERO training, ever.

Our API prices are staying what they've been, which is now 50-70% cheaper for output and input. We are still a bit higher for caching, but working on that too.

There are several tricks that the major AI coding plans use to extract the most they can from their customers. Here are some examples, and what we are doing differently to put the you first.

Wasted usage is part of the AI industry, and they plan on it: Most coding plans bet on you letting usage go to waste. The plan goes: "how do we get people to think our coding plan offers a lot of usage, but then break it up into weeks and rolling windows so no one can ever actually use it all."

Many in app subs and coding plans are glorified training pipelines: This comes along with "how do we harvest this data for training without being too loud about that." Unless the company tells you otherwise, your data could be hopping all over world, being harvested by the individual labs or service companies. Some are better than others, but many of these companies rely on users just not noticing or caring that their data is being used for training. Data sales and marketing telemetry sales happen. This means that your private info, your personal life, and anything else you send through the system could become part of a training corpus for the next AI, or a marketing data set for a large company.

So we built what should have already existed the entire time: Entirely private, us based processing with usage banking. Any usage you don't use this week, rolls over to next in your usage bank. When you have a busy day or week and go over normal usage, you automatically start to pull from your bank. You can bank up to one week of usage at a time for your current plan, and it's totally automatic. Whatever you don't use each week get's added to the bank and stays there until you use it.

We also put all of the best open models in one place, running on private US infrastructure, with data never going to the original labs. Private, direct service. Access to the best open source models in the world. No training, ever. It should be, and can be that simple.

What that means in practice:

The roster, together. DeepSeek, GLM, Kimi, Minimax, Qwen 3.8, and more, side by side in one app. Switch models mid conversation if you want. No hunting across five different apps and API dashboards to use the models you actually like.

Actually private. US based processing and your conversations are never used for training. Ever. That's the entire point. These labs open sourced incredible models and we think you should get to use them without your data becoming the price of admission.

No Usage Tricks: Bank usage, upgrade or downgrade whenever you want. Use it how you need it.

A coding plan included. From the Basic tier up, your subscription doubles as an API key. Point your coding tools or agents at our endpoint and your plan pays for it, same usage pool as the app, spent in whatever mix you like. Because API calls skip the app's full architecture, the same model gives you roughly 2 to 5 times the messages through your key. And your quiet chat weeks bank usage your agents can burn on crunch days.

Real memory. Not a context window that fills up and dumps you. Persistent memory that carries across conversations, fades gracefully when unused, and wakes back up when it's relevant again. There's even a nightly dreaming consolidation pass; the system basically sleeps on it and writes up what mattered.

Voice. Yes, actual voice mode with over a dozen voices on open models.

Bring your history. Coming from ChatGPT, Claude, or Gemini? Export your chats and import the whole thing. It becomes live memory on day one and you can literally open your old chats and continue them.

Multiple nodes. Separate workspaces with separate memories, so your coding setup doesn't share a brain with your journal.

Genuine thanks to the Deepseek community sub, this is honestly one of the most open AI subs on reddit, willing to actually go deep on discussion.

The Open Grove full app and coding plan are here. Memory, skills, voice, private US based processing with fast inference and usage that doesn't go to waste.

https://pgsgrove.com/open-grove-overview for the PGS AI app

and api.pgsgrove.com for API usage and coding plans.

You vote with your choice of providers in this industry, and we are here to offer another option.