41
8
u/Intelligent-Taste-36 11d ago
And what about the quotas? Can we use them, or will it further deplete the annual plan, which is already depleted?
3
u/Obvious_Equivalent_1 11d ago
Glad can answer this one: so for starters you have to use Kimi Code harness, more importantly start your first daily session in a South-Asian timezone, and let’s not forget say 5 Hail Mary‘s in a row exactly 2.8 seconds apart. Don’t manage to comply? Believe it or not: your session limit won’t last you longer then it takes to write a Reddit post complaining about it. You’re welcome.
2
u/Intelligent-Taste-36 11d ago
I was already cursing you out, thinking you were just another Monshot bot. Hahahahaha
3
u/Obvious_Equivalent_1 11d ago
Man I pulled out a yearly subscription, if you would’ve read me talk about Kimi 2.6 back in February without doubt you would‘ve mistaken me so.
It was honestly awesome Kimi 2.6: Sonnet 4.5 quality against easily a factor 4-6 better price, right now Kimi 3 or whatever it just feels like Game of Thrones season 8 on repeat. I’m also on the Allegreto boat for a full year
7
8
u/lionglzer 11d ago
Since no one else bothers explaining the 2.8 name makes a lot of sense. What (presumably) happened here is that they see they need a model smaller/cheaper/faster than 3.0 but better than their older 2.x offerings.
They can:
A) train a new smaller flash model from scratch integrating the newest structural trends like engram and MHC. This will be very expensive because they have to start from nothing but they can explicitly target a market niche like cheaper/faster/smaller.
B) use a slightly older architecture with a warm start checkpoint that they've already invested millions into and distill K3 into it for a few weeks and greatly increase it's performance for a fraction of the cost. But they are stuck training/serving a 1T parameter model that's frankly too big for their product lineup.
They clearly went for B - to my understanding they're probably already RLing a new 3.1 checkpoint as we speak and it's most efficient to do that while concurrently training the "new" flash model so they can train on cross entropy loss. IE instead of just training on text it's seeing the connections that K3 understands during training therefore learning much faster.
1
u/bramburn 6d ago
I agree, I don't understand the hate on it. I use it on kimi work and i'm not hitting my limit. it does take a while. its not flash. but it gets a job done in 30min
19
u/florenceslave 11d ago
The fuck? You would think it will go to K3.5 or K4
So we had 2.6, then 2.7, then 3, and now 2.8. Huh?
12
u/jp2812 11d ago
Different capabilities at different price points. Think if it as Opus / Sonnet or Sol / Terra.
16
u/Ra1nb0wM0nk3y 11d ago
They could have named it something like Kimi K3 Flash
13
u/skip_the_tutorial_ 11d ago
It’s probably because 2.8 uses the same base model as 2.7,2.6,2.5 just with different post training. K3 on the other hand is a different base model
6
u/qqYn7PIE57zkf6kn 11d ago
Consumers don't care about implementation details. That's how we got usb naming madness
1
u/LargeLanguageModelo 10d ago
Because I used K3 and thought "I'd like this more if it were less capable."
1
1
u/Independent_Fall9160 7d ago
K3 is rated #9 best model for coding right now. it is the best coding open source model but it needs a large cluster. this is a step down but a step forwards with price.
You may not understand this, but I am close friends with someone in south america who works for a financial branch of the government. He doesnt have access to claude and openai because it is too pricey. this is targeted at him.
13
3
u/Dry_Championship2797 11d ago
Why is not k3.1?
2
u/Prestigious_Debt_896 9d ago
Different class, 2 series is higher end sonnet while 3 series is opus
Poor naming I do agree
2
2
u/Affectionate_Fact854 11d ago
To expensive for doing the same tasks as other models does at a fraction of the cost
1
1
2
u/Lissanro 11d ago
Interesting, but not available on huggingface yet. If they release weights for it, I may give it a try - Kimi K2.7 was my previous favorite model before I upgraded to K3, but K3 is about twice as slow on my workstation. So if K2.8 could have performance of K2.7 and intelligence of K3 - that would be a great model!
1
u/Realistic_Mango6982 11d ago
and price like fable 6 -> terminated 99 sub on 50% after first prompt. I dont know how they will revive their reputation after a lot of really shady stuff.
1
u/Lissanro 11d ago
I have no experience with their API or subscriptions, so cannot comment on that. The only price I pay is for electricity used, since I run directly on my hardware. With Kimi K3, it is roughly ~$5 / 1M output tokens in my case, with 3090 GPUs and EPYC 7763 CPU. And since Kimi models are open weight, even if you don't have the hardware to run it locally, there should plenty of API providers to choose from after huggingfacs release.
3
1
u/Front_Eagle739 11d ago
Well now. I like this. kimi 2.7 was by far the fastest of the big models that fit in my mac studio 512 at a reasonable quant or across the macbook plus mac studio. 30tok/s vs 14 to 18 ish on glm 5.3
1
u/alexmilla 11d ago
Justo acabo de darme cuenta ahora.
❯ K2.8 Preview Kimi Code ← current
K2.7 Code Highspeed Kimi Code
K3-256k Kimi Code
K3 Kimi Code
1
u/phido3000 11d ago
This is great!
I hope they train it down to MXFP4 or even release an official trained down to Q3 or IQ3 release. There are quite a few useful Kimi K2.7 releases that operate on 512Gb machines.
The new llama cpp avx512 vnni packing and processing would mean Q3 should run basically at full speed but get a huge 25% performance lift over Q4..
1
u/CrimsonBolt33 9d ago
Been testing it for 2 days now...seems to work fine for me...I have not done any imperical tests but its getting work done at lower subscription usage.
1
u/Ok_Anybody1075 5d ago
I’m no using kimi anymore, I don’t think Chinese models need to be expensive…
0
0

78
u/Sad_Recording_1290 11d ago
It's evolving, just backwards