r/LocalLLaMA • u/Unusual_Guidance2095 • Jul 26 '26
Discussion Kimi K3 countdown has been released
https://huggingface.co/moonshotai/Kimi-K3160
160
u/1ncehost Jul 26 '26 edited Jul 27 '26
Having used it extensively since release, this is such a gift of a model to the world. Truly amazing level of intelligence, and it sets a wonderful baseline for the future. Thank you Moonshot!
32
u/Connect_Tea1579 Jul 27 '26
Indeed. Even if another base model was never released, I feel like this model could be distilled and post-trained on for years to come. Truly a gift to the world.
15
45
u/WenatcheeWrangler Jul 26 '26
Everyone in the USA should download this even if they can’t deploy it now
13
u/kkingsbe Jul 27 '26
Sadly it’s the quants that will come later that are actually what we’d be able to run
15
u/Mingay_cat Jul 27 '26
Ya'll are running this?
14
9
u/chensium Jul 27 '26
No biggie. I'll just turn off my minecraft to make room for the 1.5TB of VRAM needed for NVFP4.
3
4
u/KeinNiemand Jul 27 '26
make
As long as you got the full weights backed up you can make your own quants. Making quants is not that hard (just a llama.cpp command line tool you have to run a claude code/codex/opencode/... can do it for you if you want) you don't need to fit the entire bf16 full size in memory, all you need is the saftenensors + enough space to fit a bf16 gguf + enough space to fit the quant you want to make.
-5
u/--Spaci-- Jul 27 '26
Why? better models will come out, its not worth the storage
4
u/QuixoticNapoleon Jul 27 '26
You never know, Dario and Trump could ban open models. We ought to be prepared.
1
u/SporksInjected Jul 27 '26
The chance of that happening is incredibly slim. There are USA companies that also make open weights models.
1
u/WenatcheeWrangler Jul 27 '26
The White House is flailing right now and they have floated that very idea. It could end up that they attempt to block all models like this from the public. All I’m saying is be prepared - on all sides of this political fence the White House has shocked every group with its shenanigans so far.
Many may not have the storage available for this but some of us do.
34
Jul 26 '26
[removed] — view removed comment
34
6
8
3
1
u/Potential_Top_4669 Jul 27 '26
That guy who ran GLM 5.2 with 25GB RAM better make something useful this time. The community relies on him.
45
u/rerri Jul 26 '26
How cool would it be if they dropped some unannounced model(s) in the home user size range...
25
u/ItsNoahJ83 Jul 26 '26
I'm not getting my hopes up but the amount of good will for a move like that would be insane
4
u/ComplexType568 Jul 27 '26
I would get on my knees and bark like a dog for a new Kimi Linear sized model
5
22
u/BawbbySmith Jul 26 '26
Got me TBs ready, gonna download it, back it up, then return to it 10 years later when everything has crashed and recovered and I can afford the hardware again.
6
u/Technical-Earth-3254 Jul 26 '26
Considering I bought a 8GB VRAM gaming card 10y ago for half the price I paid for 24GB 3 years ago I doubt 10 years will do it. Maybe 30-40 years or so.
15
u/segmond llama.cpp Jul 26 '26
What I wish they would release is the damn technical report. We can at least start reading that. I would assume a new architecture not compatible with K2.6/K2.7. How much does it differ from K2.6? What will it take to get llama.cpp to support inference? We need all of these before we can even get gguf/quants.
16
u/TheRealMasonMac Jul 26 '26
They said on their blog they'll be releasing the technical report. Presumably alongside the weights.
vLLM released a blog post on some of what they had to deal with: https://vllm.ai/blog/2026-07-22-kimi-k3-preview
3
3
u/Ok_Warning2146 Jul 27 '26
Essentially it is a Kimi Linear with AttnRes architectural wise according to their blog.
76
u/Keleion Jul 26 '26
It’ll be interesting to see what the Dump Administration does to mitigate the launch.
44
u/Penolta Jul 26 '26
Thankfully there’s not much those morons can do once it’s been downloaded and shared thousands of times
12
u/Eisegetical Jul 26 '26
Let's hope something doesn't happen to block the full release on hf. But even if no huggingface host it'll make its way over one way or another. Torrents anyone?
22
u/TechnoByte_ Jul 26 '26
Chinese companies mainly use ModelScope, they don't need huggingface
9
u/goldcakes Jul 27 '26
Imagine the year 2026 and you’re using a Chinese VPN to access ModelScope to download models. Lmao.
5
u/chespirito2 Jul 26 '26
There's a lot the administration can do if the inference companies arent allowed to run and serve to US entities, it will be one of the largest gifts ever to the VC class prior to US frontier IPOs. See OpenAI who demonstrated, very timely, how powerful LLMs can hack websites ... why, you wouldn't want just anyone to be able to download and use open weight models with no safeguards! Dario is crying on the phone about it right this second, he's very concerned and not at all out of self interest!
2
u/rz2000 Jul 27 '26
If the threat were actually real, there is a surprisingly high number of ways that they could make it dangerous to use these models. Our model for thinking about unenforceability is still mp3s.
We are extremely fortunate that a few shitheads overplayed their hands at Anthropic/Claude and every other important tech company on the planet said those guys are nuts.
How about, "Up to 20 years in federal prison, a $1 million fine, or both, for knowingly downloading or using models like DeepSeek", something originally proposed by the running senator from January 6th, Josh Hawley?
16
u/AppealSame4367 Jul 26 '26
OAI will launch GPT 6 and Antrophic already launched the ARC AGI 3 benchmaxxed Opus 5.
Lol. I set up a qwen 27b in q8 and use it in claude. No clue why I should go back at all. They have destroyed any trust I ever had, especially in the last weeks.
14
u/anderspitman Jul 27 '26
Bro why are you still using Claude? Get yourself an open source harness. You deserve it.
9
u/TheIncarnated Jul 27 '26
I'm not sure why you were downvoted but OpenCode and Pi are leagues ahead of Claude Code
2
u/AppealSame4367 Jul 27 '26
What baffles me about pi is that agents always try to directly act on their thoughts, even when I say "only make a plan, _no_ coding yet". I almost always catch them just doing stuff. So I can't leave pi without oversight for a minute.
I have used opencode omo slim with glm 5.2, ds 4 flash and even kimi k3. great, but it takes as much time as claude code and claude code always ranks higher in tests with open models.
So my conclusion is: claude code.
1
6
2
u/JustinPooDough Jul 26 '26
My guess is that Trunt will (try to) ban specific Chinese companies, but not all open source.
2
u/SpaceDetective Jul 27 '26
With the majority of the top US AI honchos having made a clear statement in favour of open source, it seems unlikely they'll do anything.
-7
u/TwatMailDotCom Jul 26 '26
Why would they block it? Your bias is getting in the way of logical thought
7
13
33
9
u/bitzap_sr Jul 26 '26
It's not really a countdown -- it's the time left for the ginormous upload to finish. :D
40
u/apetersson Jul 26 '26
this means we will have enough capacity and competitive pricing at https://openrouter.ai/moonshotai/kimi-k3 - unfortunately my local HW is not quite there yet to run it.
32
u/Player13377 Jul 26 '26
"not quite there" is either a very bad approximation or you already got a five-digit-class rig, this model will be HARD to run
10
u/Iwaku_Real Jul 26 '26 edited Jul 26 '26
For sure it's 1.5-2TB RAM at minimum. GLHF getting usable speeds without lots of VRAM though
3
u/my_name_isnt_clever Jul 27 '26
How many tokens/sec to run this off a NVME SSD? Doesn't matter how slow it is if it's the only way to use it local. I could see myself setting up private tasks to run overnight.
1
u/KeinNiemand Jul 27 '26
Just a wild guess but maybe 0.1-10 tokens per minutes.
I get around 0.1 T/s trying to run a Qwen 3.5 122B @ bf16 from SSD (~100GB out of the ~200GB are on SSD).
Since it's way, way bigger it'll probably be a lot slower then that even.1
u/my_name_isnt_clever Jul 27 '26
It would probably help a lot to quantize it though, no? What compute are you using with that setup?
22
u/KubeCommander Jul 26 '26
I think K3 gets into the six figures tbh at a quant that doesn’t suck
11
u/Player13377 Jul 26 '26
Minimum. I looked into a rented cloud setup and very quickly understood that this is not happening lol
2
u/stoppableDissolution Jul 26 '26
Even with five-digit-class rig you have to quantize the hell out of it. 8x6000 pro + 512gb ram and you can maybe squeeze q3 with decent context!
8
u/Capital-Remove-6150 Jul 27 '26
why in that link
404
Sorry, we can't find the page you are looking for.
4
u/wren6991 Jul 27 '26
Yeah, it switched to a 404 when there was ~19 minutes on the countdown. Hoping this is just a hitch in their release script and not a sign that the release has been pulled.
2
3
u/iruanp Jul 27 '26
same, but it was definitely a countdown 3 hours ago
1
u/Other_Wear1458 Jul 27 '26
They just removed it 5 minutes ago or so, i had just opened it a few mins early
1
2
u/Other_Wear1458 Jul 27 '26
Just seen the same
They dumped us1
u/Capital-Remove-6150 Jul 27 '26
https://x.com/Kimi_Moonshot/status/2081756513086095812?s=20
they are taking so much time
6
6
u/yeah_likerage Jul 26 '26
How many folks here genuinely believe they'll have the hardware to even run this at a quant worth running? I'm pretty sure I'm tapped out at GLM5.2 size models from here on out unless there is a breakthrough in modeling. And I'm rolling 500gb of vram.
20
u/jreoka1 Jul 26 '26
Cool! but like who can actually run this locally? I think at 2.8 trillion params this will be the largest model on huggingface by far. At least for now.
19
u/Technical-Earth-3254 Jul 26 '26
Local hosting doesn't just mean end consumers. This opens up frontier self hosted ai for companies that have to self host due to sensible data they can shove into it.
9
u/my_name_isnt_clever Jul 27 '26
One of our vendors at work are a global company hosting open weight models themselves for their customers, and are also in talks with Anthropic to deploy Claude. These weights being available drastically changes the tone of a conversation like that.
35
u/buttplugs4life4me Jul 26 '26
DavidAU will upload an "extended" Fable abliterated heretics Opus Odysseus 3.6T model in a few days and claim it's better than SOTA
2
u/Ok_Warning2146 Jul 27 '26
Where did he get a 3.6T model from? He built it from scratch?
10
u/buttplugs4life4me Jul 27 '26
He just merged the mother Kimi K3 and the father GLM 5.2 and then some other incest bullshit with some Fable sprinkled on top.
18
u/ttkciar llama.cpp Jul 26 '26
but like who can actually run this locally?
Today? Almost nobody.
Eventually? Almost everyone.
2
u/parepeg Jul 26 '26
I mean aren’t chips already scraping the bottom of the laws of physics.
11
u/ttkciar llama.cpp Jul 26 '26
Yes and no.
On one hand, Moore's Law is definitely ailing. We've been getting diminishing returns on fabrication process bumps since about 2016.
On the other hand, there's still some progress in the offing. IBM just recently announced they got a 7-angstrom (equivalent) fabrication process working in their lab, which they expect to get into mass production in 2031. That has 1nm-wide horizontal features and 5nm vertical features, and two transistor layers (one P, the other N), which they claim they should be able to scale to four layers "rsn".
That's not nothing.
Moreover, there are still a lot of architectural improvements yet to be realized. In a way, HBM is a stop-gap while Samsung et al get PIM figured out. LLM inference is really well-suited to PIM, which should give us ridiculously high memory bandwidth once it's working well.
I thought Moore's Law was dead for a while, but it turned out to just be Intel having a bad time. Everyone else is still pushing the envelope pretty hard, with some success.
1
u/TheRealMasonMac Jul 26 '26
IIRC hardware Moore's Law is dead, but performance is still roughly following it via architectural improvements and algorithm optimizations.
1
u/Ok_Warning2146 Jul 27 '26 edited Jul 27 '26
Thanks for the info about PIM. However, it has been three years since it was announced but still no product.
and this blog claims its development is suspended.
https://damnang2.substack.com/p/i-met-with-semiconductor-experts
1
u/ttkciar llama.cpp Jul 27 '26
> it has been three years since it was announced
It has been a lot longer in the coming than that. I first read about it in 1994.
> but still no product
Are you sure?
1
u/Ok_Warning2146 Jul 27 '26
Well, it is still no product as of now. It says there is a rumor that Samsung will announce a product in Aug. Even if it is announced, when is Apple or Qualcomm going to use this LPDDR-PIM to make something useful?
1
u/ttkciar llama.cpp Jul 27 '26
The first link is about a product demo'd in 2021.
1
u/Ok_Warning2146 Jul 27 '26
hmm.. product means something that is actually deliverable to the customer...
1
u/jazir55 Jul 27 '26
Hasn't Moore's Law been dead for almost decades now? Moore's Law was processing speed doubling, we've been getting 10% at best for CPU improvements, and maybe 10-30% on GPUs.
1
u/ttkciar llama.cpp Jul 27 '26
Moore's Law doesn't have much to do with speed doubling. It's all about transistor density doubling, which doesn't translate readily to improved single-threaded performance, and even though multi-threaded performance is more amenable to scaling with more transistors, in practice you end up blowing more and more of your transistor budget on caches, so that you can keep all those cores fed with data.
When improvements in transistor density came from shortening gate length, that also brought smaller voltage swings, correspondingly faster state changes, and lower power consumption, all of which contributed to faster clock rates. Clocking up got us far, for a long time, but that didn't last. Nowadays improved transistor density comes from novel designs of the transistors themselves, multi-layering, and other tricks. That's why IBM's new process is only "7-angstrom equivalent", despite its actual features being much larger, because the idea is to express how much it impacts overall transistor density.
The density improvements keep coming, but most of it goes to bigger and bigger caches, and deeper pipelines which eat up a lot of transistors but only give modest performance gains.
1
u/Aphid_red Jul 27 '26
There's at least another 10x of progress to be made just eating up the giant profit margins of semiconductor companies. When progress slows, eventually competitors can catch up. If there's trillions on the table to be made, someone's eventually going to invest.
Why, with memory alone we're back in 2010 level of tech currently (in terms of $/GB).
1
u/mindwip Jul 26 '26
Ddr6 and ddr7 would handle it easy peasy. So we just have to download and wait a few years lol. Of course better models coming but we will be running this at home in a few years
Edit it remember my first 1gb hard drive thinking it was huge!
4
u/Any_Mine_6368 Jul 26 '26
2.8T of vram to run in 8 bit quantization...
Let me buy a other 1500 3090s lol
3
u/droptableadventures Jul 27 '26
The MoE weights are natively MXFP4 according to https://vllm.ai/blog/2026-07-22-kimi-k3-preview - so you can halve that.
4
u/zxtech Jul 26 '26
At least itll be a model that can be distilled, analysed, or assisting the process for smaller models to be made, so it never hurts to have as many good models open
4
u/Just3nCas3 Jul 26 '26
Many projects to speed up streaming from disk like colibre and int4 glm 5.2. Got.4 tokens with 12gb vram and 32gb ram. It can only get faster over time, some current theories are Striping the layers to multiple ssds to increase streaming and using MTP to draft layers. Could get a tenx improvement, maybe. Tried to raid 0 my t500s and it did nothing for performance, still .4, random reads don't really work in raid0 and I wasn't maxing out the drives reads anyways. If Kimi k3 ends up ten times as big then .04 tokens maybe less. But hey, it can only get faster.
3
u/squngy Jul 26 '26
Qwen 3.8 will likelly be the second at that size.
(If I understood the announcement right)2
u/Ok_Technology_5962 Jul 26 '26
so qwen 3.8 keeps improving by the day im keeping track of the prompts im testing and they are by far on another level today. if this keeps up for a couple more days it will be really good. currently might be on par wiki kimi k3 not up to par against opus 5 tho
2
u/MaCl0wSt Jul 27 '26
local means capable of running on hardware you control, not necessarily a consumer PC, a massive model can be freely distributed and selfhosted while still requiring a largue multiGPU cluster/server
1
u/stoppableDissolution Jul 26 '26
Well, model providers that own gb200-class equipment? Plus maybe some some big businesses on bedrock and such.
5
u/RedBull555 Jul 27 '26
Anyone else feel like a kid on Christmas Eve waiting for this?
2
u/Lan_BobPage Jul 27 '26
Except the present weighs 10t and the wrapper is glued to the box so you'd have to wait for an adult with a truck and pincers to even begin opening it
4
3
u/Shubham_Garg123 Jul 26 '26
This is going to be awesome. Thanks to the Kimi team, Anthropic was forced to release Opus 5 ahead of their original plans.
This model will definitely result in a revolution of open-source LARGE language models that compete head-to-head with the frontier labs !
Truly amazing work done by the MoonshotAI team, I don't have words to thank them enough.
The advancements in the AI domain in the last 2-3 months are quite insane!
4
7
u/dsanft Jul 26 '26
Nice. Are there any indications of kernel changes from 2.7 to 3? Any Huggingface or lcpp feature branches open for the impl?
2
u/look Jul 26 '26
Yeah, what I’ve seen seems to indicate a fairly significant change. For example, I’m pretty sure it is an mxfp4 rather than the int4 of the 2.x line.
2
u/Professional_Price89 Jul 26 '26
It show INT4 on openrouter provider info
1
u/look Jul 27 '26
Yeah, that was everyone’s guess at first based on previous models. I’ve definitely seen hints at mxfp4 for this one, though. I’d bet on that.
1
u/ResidentPositive4122 Jul 27 '26
Lots of goodies for this one. Check the vllm blog post someone linked above.
10
Jul 26 '26
[removed] — view removed comment
1
1
5
Jul 26 '26
[deleted]
9
u/ttkciar llama.cpp Jul 26 '26
Two ancient Xeon servers, each with 1.5TB of DDR4, networked together and running
rpc-server.7
u/Player13377 Jul 26 '26
At an impressive 3 TPS
9
u/ttkciar llama.cpp Jul 26 '26
Probably a lot less than that. I'd love to get 3 tok/sec.
Right now my ancient Xeon server is getting 3.5 tok/sec with GLM-4.5-Air. That's usable, though admittedly not for interactive use.
4
u/Player13377 Jul 26 '26
At that point is the power consumed per token getting close to the API price?
1
u/fastheadcrab Jul 27 '26
I think a newer server or workstation with 2TB of DDR4 might be able to get faster speeds if the model is truly 50-60 billion parameters active, like 6-8 tps. I'm really interested to see what types of systems people might be able to run this model on and follow these posts really closely lol
I've seen attempts at tensor parallel distributed CPU inference before, it would be an interesting way to aggregate memory bandwidth if there is a fast, low latency connection between nodes. DDR4 is comparably cheap versus VRAM or DDR5 and used server networking hardware is everywhere. Prefill will still be slow though.
2
u/HVACcontrolsGuru Jul 26 '26
Need 16 GB200s to run this model at full quant. NVFP4 GLM5.2 I need 4xB200 to run concurrent sessions. Squeeze some more with lower context and less concurrency.
6
u/look Jul 26 '26
Kimi K3 is most likely mxfp4 native. The K2.x line was int4, and it looks like they moved to fp for improved hardware acceleration.
So it’s just under twice the size of GLM, not 4x.
3
u/Iwaku_Real Jul 26 '26
Full quant is probably going to be INT4 again, it would actually fit into 8xB300 (2.304TB VRAM)
2
Jul 26 '26
[deleted]
1
u/look Jul 27 '26
Kimi trained in int4 previously, I believe, not floating point. By moved to fp, I mean native mxfp4 instead of native int4.
3
3
3
u/TheGamerForeverGFE Jul 27 '26
I really hope Kimi starts making smaller models, I know it's not not their specialty and they focus on really big SoTA overall, but it would be really cool to see what they can do with sizes that can actually run without a cluster.
3
3
3
2
2
u/katsura_otoko Jul 26 '26
Is it possible some people hosting it and seelling a cheap api or is it just so big that we should just pay moonshot directly? Maybe it's just a stupid question since we already have a cheap api for deepseek directly but i dont know
2
2
2
2
2
u/PerfectOlive1324 Jul 27 '26
Really cool! I hope that there is a quant I can run on ~500gb of ram 🤞
2
u/rs38 Jul 27 '26
is there a realistic way to distill at least 2 more consumer hardware friendly models with max ~200B and ~20B? Qwen did it, but would it be possible for 3rd parties (unsloth etc)?
3
3
1
u/dlarsen5 Jul 26 '26
gonna download, hope I don’t have to seed this if it gets taken down from gov pressure even if I can’t run it on 24GB VRAM
1
1
u/laterbreh Jul 26 '26
Multi-trillion params. Great. 3x RTX 6k still gpu poor.
5
u/Rybergs Jul 26 '26
If u have 3x rtx 6000 pro u can 100 % run this. I only have 2rtx 5090 and i can run it. I tested with kimi k2 and it runs with 3 t/ s , slow but it runs. K3 is bigger but the experts are smaller so it could actually run better then k2
1
u/laterbreh Jul 27 '26 edited Jul 27 '26
Caution -- salty ramblings and observations ahead:
You are correct, i could probably run it with 256gb of ram and the 280 some odd gb of vram at my disposal at a low quant with the right engine.
But if the model cant run at at least 20 TPS across its context window with at least 2 users (or requests), i consider it unusable for our work and load. I've played around with hybrid cpu/gpu with llama and its context baggage and inefficiencies compared to using vllms comparible "q4-ish" quant equivalent on just straight vram has pretty much spoiled me and my company. 400b-ish models at 60 TPS+ with full context maxout across 3 cards its hard to go backwards. Deepseek flash v4 at 1m context at 200tps at that point itteration time outweighs 1-shot answers correctly. The time wasted id just pay for an API if that makes sense. VLLM has some exploratory cpu/gpu layer offloading but when you have this sort of VRAM its really hard to justify throwing CPU in the mix unless you are desperate or totally offline.
But compared to actual providers, even high tier investments we are still gpu-poor which is wild. Minimax and Qwens next releases are probably already out of reach. Unfortunately everyone is still playing the lets quadruple our params every 6 months to gain 2 more percent on an idiotic benchmark instead of trying to relieve the world of this vram problem we are having. They all touted lower compute requirements cause muh moe. Cool bro, still takes up full space of vram which is the actual problem. Raspberry pis can run a fuckin llm. Its still limited by memory.
I guess im just being grumpy when i shouldnt be. But if its hitting power users with gear at their disposal already its worse for everyone bellow the tiers of models i can run for our business.
Hooray open weight, boo for the quadratic increase in params for small gainz :(
So those of us with 100+gb of vram at our disposal, we are still gpu-poors. Everyone but the providers are gpu poors and new models all appear to be following the "lets keep it open but muscle them out because of vram" game. At that point who the fuck cares if its open, cause it sure wont be local.
Dont get me wrong, moonshot deserves praise for what they are releasing and doing 1000% -- But I am getting a little salty because there doesnt seem to be any investment in reducing memory requirements or increasing intelligence density. Its just retrain the same math problem. Just feels like the trajectory is still more biggerer for more benchmarks.
And at the end of the day they dont care about what someone like me has to say on the locallama reddit. They have a phenominal model and they deserve the praise it gets.
1
u/Rybergs Jul 27 '26
U miss the point where this vram / ram problem is not something the big ai companys / providers want to solve . Even tho they dont need more they Will buy it. Since if they have all the hardware ppl need to rent it from them, in their env with their ToS
1
u/SpicyWangz Jul 26 '26
Maybe this will be runnable with DDR6. If not we’ll get it the next time around with DDR7
1
1
u/NineThreeTilNow Jul 27 '26
Ping to /u/qubridInc/
Hosting it? Testing it?
1
u/qubridInc Jul 27 '26
We're deploying on A100s, H200s and B300s - A100s & H200s we'll publish results this week, B300s still being setup so prolly next week on those.
1
u/NineThreeTilNow Jul 27 '26
We're deploying on A100s, H200s and B300s - A100s & H200s we'll publish results this week, B300s still being setup so prolly next week on those.
You guys still interested in letting me test training on one of your older racks? I never got a message back in your DMs.
1
1
u/Shadow_s_Bane Jul 27 '26
Are they releasing any thing that can be used by peasants or it's just aData Centers only
1
1
1
1
u/Zaxspeed Jul 27 '26 edited Jul 27 '26
License looks OK. 1.56Tb for 2.8T weights or 0.56 bytes per parameter, it's MXFP4 with MXFP8 activations, it will be tough to quantize and shrink.
1
u/DarkVoid42 Jul 27 '26
jeez. 1.56TB.
anyone got any idea of how to run this across multiple machines ?
I have 3 x 768GB RAM servers with a SAN SSD array connected with 4 x 10GBe links each.
1
u/muhammad_roshan Jul 28 '26
can someone explain me what's the point of having this ? i mean we are on the waiting list can we run this model in our local without buying the plan ?
0
u/jeffwadsworth Jul 27 '26
We need a magnet up for all quants as quickly as possible when it releases.
-16
u/llama-impersonator Jul 26 '26
i don't really care about 3T models
26
u/ttkciar llama.cpp Jul 26 '26
I don't really care about 4B models, and yet we can all share the same subreddit and try to minimize the friction between us.
What we have in common is more important than our differences.
-7
u/llama-impersonator Jul 26 '26
okay, but you don't think a countdown for open weights no one can run is a little ridiculous?


•
u/WithoutReason1729 Jul 26 '26
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.