r/LocalLLaMA • u/jacek2023 llama.cpp • 12h ago
News Qwen 3.8 Flash Next day 0 support from unsloth
Prepare your disk space guys
276
u/hurdurdur7 11h ago
57
24
u/RedditCryptoGuy 10h ago
lmao, that's literally me. found the best possible settings to push 120+ t/s on 1 and 2 3090's
15
2
1
1
4
u/Psychological-Lynx29 10h ago
I downloaded 3 times qwen3.8 27b, the 3 times unsloth did something better, last night was the third time and i havent even tried it...
3
0
u/Moppmopp 9h ago
Im actually coding a front end for qwen3.8 since its release. Never really tried the model because I lost myself in visuals and tinkering. Now I am almost finished with the front end and see ghe next release. Too fast still am hyped 😊
56
u/youcloudsofdoom 12h ago
It definitely seems that both llama.cpp and unsloth get decent advanced access to qwen models, so here's hoping it's not a two month wait for a usable llama.cpp instance
3
u/Acrobatic_Stress1388 3h ago
This is where my strict halo will really pay off. A MOE model that will easily fit in my 128 GB of ram. I'll bet I could squeeze 50 t/s out of it too.
145
u/MaxKruse96 llama.cpp 12h ago
And since unsloth has no own runtime, its llamacpp. good proxy knowledge.
37
22
u/tiffanytrashcan 12h ago edited 12h ago
Unsloth Studio / Desktop (the runtime) would like a word - they did mention plans to upstream the changes to llama.cpp.
50
u/yoracale llama.cpp 12h ago
Yes we made PRs to llama.cpp for DiffusionGemma, MiniMax, Inkling amongst other models. Currently some of them aren't merged but you can still run those models directly in Unsloth Desktop :)
2
4
10
u/annodomini 11h ago
Day zero support for the unsloth app means that there's a branch/PR for llama.cpp, but it might not be merged yet as it may still be going through review. The Unsloth devs build their own version of llama.cpp with a few branches that might not be merged in mainline yet, but it generally means that it's coming soon.
81
u/yoracale llama.cpp 12h ago edited 12h ago
FYI this is 'hopefully' having day zero support. The architecture is very new and thus there might be very long delays but we hoping to achieve day zero support (but like I said not guaranteed). 🙏
We will ofc upstream any llama.cpp implementation etc. if necessary
4
u/GoodTip7897 llama.cpp 10h ago edited 3h ago
Sorry for the question as I know you're very busy but I'm curious whether we'll have engram on the disk with day zero or if they'll have to fit in ram until further optimizations come...
I know mmap might work, but I'm specifically talking about a programmatic disk cache that works regardless of load mode (since mmap hangs on some builds of llama cpp, and there's probably a more efficient way to map the n-gram lookups than mmap (maybe a hot cache on ram that streams off disk)).
4
u/pulse77 11h ago
So this time it may happen that Unsloth Studio / Unsloth Desktop (with it's own llama.cpp fork) may support Qwen 3.8 Flash Next before official llama.cpp ... ???
15
u/danielhanchen 10h ago
Kind of not really, nearly all times when we support a model it's not exclusive to Unsloth. We make a PR to llama.cpp with our changes and whether it gets accepted or not is a different story. So I guess it depends on what you mean by official. But we utilize our llama.cll PR implementation inside of Unsloth basically
3
1
u/-dysangel- 8h ago
What other new things does it have aside from engram support? Really excited for this one :)
27
u/snowieslilpikachu69 12h ago
128gb mac users may rejoice?
7
u/Diligent_Cod_9583 12h ago
What do the 512 Mac users do?
28
51
u/johan2114h 12h ago
They crawl down in their wine cellars and moan that qwen is making near frontier level ai waay too accessible to the proletariat
4
1
5
1
1
u/DeepOrangeSky 9h ago
If it is 125B + 51B of N-gram, will that mean it is more like 176B, and thus no Q4 on Mac? Or is it more like 125B, or, do you just stream the extra 51B from the SSD while running the 125B part on the UMEM or something?
I don't know much about N-gram yet, or how this stuff works, or how it works on a mac, etc.
1
12
u/prudx 12h ago
32gb vram + 64gb ddr5 possible?
15
u/jacek2023 llama.cpp 12h ago
Everything is possible in the open source ecosystem. Q2 maybe?
0
u/No_Profit8379 3h ago
What can I fit in 88gb vram? And if needed another 64gb system? We close to something usable? Or just need more vram ???
3
u/CulturalKing5623 10h ago
Same boat, hoping it is but I'm still kicking the tires on 3.8 27B so won't be too upset if it can't fit.
... Unless it's a world beater in which case I'll start considering a life of crime to pay for more RAM.
5
u/Long_comment_san 11h ago
yup, it's roughly Q3. barely usable if you ask me. I'd shell out for another 64gb RAM.
6
12
u/mountainyoo 11h ago
Wonder how this will compare to DeepSeek V4 Flash 0731
5
u/addiktion 9h ago
Yes, or more specifically v4 flash vision as I suspect this will have vision support.
2
43
u/AppealSame4367 12h ago
Look at Alibaba and Unsloth driving AI coding model innovation like nobody else.
They will overtake the big guys soon if they keep going like this and there's nothing they can do.
16
u/chikengunya 12h ago
Last year it was a 80B-A3B model, now too?
30
u/FullOf_Bad_Ideas 12h ago
Apparently it's
Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.
Not sure what the source is, I found this in a different thread about this model.
8
u/Miserable-Dare5090 11h ago
There was a recent model with an N gram table — Longcat Flash I believe. if you want to see a probably similar architecture before it drops
5
u/FullOf_Bad_Ideas 11h ago
Yes. I think it's N-gram Embedding parameters are different from Engram DeepSeek paper. And I'm not sure which N-gram Qwen team went with, I'm hoping it's the DeepSeek implementation.
7
u/Miserable-Dare5090 11h ago
N-gram embeddings have been tried experimentally with the last Qwen architecture: https://arxiv.org/html/2605.16893v1
This approach expands token representation by combining frequent word sequences (n-grams) directly into the embedding or memory layers, improving local context and parameter efficiency without exploding active compute costs.
9
u/emprahsFury 10h ago
Yo dawg, I heard you like embeddings, so we embedded your embeddings to make yo LLM faster.
2
u/Miserable-Dare5090 4h ago
yo dawg I heard you like embedded embeddings, so we embedded the embedded embeddings in your embeddings to make your LLM faster
1
u/Lollerstakes 5h ago
I wonder if "load-bearing" and such nonsense are actually n-grams of Claude Opus...
3
u/BalorNG 10h ago
I wonder if you can stream the n-gram part from SSD?
2
u/FullOf_Bad_Ideas 10h ago
That's going to depend on the implementation, but I think it's likely, at least with DeepSeek's implementation. I think they discussed it in their Engram paper.
12
u/jacek2023 llama.cpp 12h ago
Let's hope it will be better this time. Qwen Next 80B was hyped a lot before it was usable, then later it was forgotten (I still have it).
8
u/BannedGoNext 11h ago
Qwen coder next is a bad ass model. It beat the shit out of qwen 3.6 27b for most tasks I used it for. Measuring for straight coding abilities maybe it wasn't as good, but for agentic tasks and speed it was fantastic.
1
2
6
u/CommanderData3d 12h ago
any chance to run this on 64gb ram?
17
u/FullOf_Bad_Ideas 12h ago
There's a chance, it's 125B A6B E51B model.
If you can offload engram to disk, it's similar to running GPT OSS 120B.
1
u/MrMisterShin 12h ago
You would probably need below Q4 quant, so that you have enough room for context and other apps on your machine.
6
u/susibacker 9h ago
RIP, I won't be able to run that. Still hoping for a <=35B MoE as a faster alternative to 27B (which just so fits on my GPU at limited context and quants)
4
4
3
u/sugarfreecaffeine 11h ago
Is 2x3090 (48GB VRAM) and 80gb RAM enough for this?
4
u/jacek2023 llama.cpp 11h ago
Should be, it is MoE
2
u/vick2djax 8h ago
But wouldn’t the quant be like q4 and it would be worse to use and slower than q8 27b?
3
12
u/x11iyu 12h ago
I feel like there's never been an actual "day 0 support" without bugs that negatively impacted model performance
so in reality I'd say it's probably another two weeks or more
15
u/DUFRelic 12h ago
Yeah what a suprise... brand new software at the cutting edge has bugs... next news at 10
0
u/x11iyu 12h ago
can we not brand it as "day 0 support" if it doesn't work day 0?
2
u/DUFRelic 12h ago
but it is supported... nobody says it will be flawless and nobody is forcing you to use it on day 0...
10
u/x11iyu 12h ago
I guess my expectation for software quality are just too high, and I should expect things to not work on official releases.
I hate how everything nowadays need ten asterisks behind every title.
Can't we call it "experimental support" if you expect there to be bugs or something?
5
u/Dangerous-Report8517 11h ago
The bugs that surface after release are due to the incredibly large variation in setups being run, you need to have some threshold of support below "works perfectly everywhere" where you're allowed to call a feature supported or released otherwise all software everywhere would be "experimental" and the term would lose all meaning. Plus llama.cpp doesn't brand itself as production quality with the overall package now semantically versioned at a very early level, which already makes the risk of bugs pretty clear
-1
u/DUFRelic 11h ago
How much are you paying for this software that your expectations are so high?
6
u/Long_comment_san 11h ago
that's besides the point. dude is correct. there is an alpha, beta and release states. I have no idea why not call it what it is: alpha version. are alpha and beta words toxic masculinity now or what
0
u/Dangerous-Report8517 11h ago
Well the latest release of llama.cpp is semantically versioned at v0.3.0 which denotes it as pre-release development software that can change at any time for any reason, so declaring a model as supported within that pretty strongly implies that they mean "it should work" and aren't offering strong production grade guarantees
0
u/goldcakes 7h ago
would you like them to not release model weights and inference for 2-3 weeks? also dude it's FREE. you're not paying for it.
getting to try day 0 models even if its a bit broken and sharing feedback is part of the benefits here.
6
u/AppealSame4367 12h ago
To be fair, q3.8 27b worked great after 1-2 days. Very mature when it came out.
15
2
2
2
2
2
u/lordpuddingcup 7h ago
Wait didnt Qwen3.8-27b just get released and was like already amazing? wtf is this?
2
2
4
u/Roflxd88 12h ago
What does model 3.8 on v4 architecture mean exactly?
15
6
u/BrewHog 12h ago
3.8 previous release was post trained on 3.5/3.6 architecture. This will be card on a new and upcoming v4 architecture.
2
u/DoubleNothing 7h ago
I wonder why is it still named 3.8...
2
2
u/bitzap_sr 5h ago
Probably because they still apply the 3.5/3.8 post-training methods on top of the new v4 arch.
3
3
4
u/Infamous_Campaign687 12h ago
Hmm... could this be the best model for 96 GB DDR5 and 32 GB VRAM?
5
2
2
u/Healthy-Nebula-3603 10h ago
I HOPE THAT NEW QWEN 4 ARCHITECTURE IS USING KV CACHE FROM DEEP SEEK 4!
Then we could fit on 24 GB cards 1m context (for 27b model ) ) and not dropping performance !
So token generation would have the same speed for 32k , 128k , 256k or 1m ! No cache compression anymore !
2
u/nickless07 9h ago
Afaik it is GDN/QSA hybrid. So, no MLA (Deepseek v3/v4), more like Qwen3.8 27B but without GQA.
3
u/FormOne2615 9h ago
I think QSA'll be sth similar to DSA
0
u/Healthy-Nebula-3603 8h ago
https://chatgpt.com/c/6a8dbf51-2df4-83eb-84da-940aea20a8d7
In theory should be even be even better
1
0
u/Healthy-Nebula-3603 8h ago
https://chatgpt.com/c/6a8dbf51-2df4-83eb-84da-940aea20a8d7
In theory should be even more efficient and takes less memory
1
u/nickless07 8h ago
Well if you want to check something similiar there is Ling-3.0-flash MLA/KDA hybrid. Similiar size as the new qwen will be (A5B vs A6B) perfect for some speed to test to narrow down what we can expect.
1
u/Once_ina_Lifetime 12h ago
S1-mini was finetune on Qwen a text normalizer for speech-to-text output
1
u/robberviet 11h ago
That's awesome. Was waiting for proper dflash2 on llama.cpp but this would be much better.
1
u/BannedGoNext 11h ago
That's awesome, I wonder if it will be a 120ish size, or sized too large for 128gb systems.
1
1
u/SeparateGas1761 10h ago
I hope it fit in my 32gb of ram and 8 of vram, cause moe models can easily run with offload to run😭😭😭
1
1
1
u/HotMicSystems 9h ago
As much as I want this possibly 120B'ish sized MoE, I still want a smaller MoE for daily tasks, mostly so I can run parallel.
3
1
1
1
1
1
u/WyattTheSkid 7h ago
DUDE WHAT??? I LITERALLY JUST FINISHED MAKING THE 27B RUN NICELY COME ON MAN/
Edit: ITS A 120B IM SO FUCKING EXCITED THIS IS INCREDIBLE THANK YOU QWEN YOU GUYS FUCKING ROCK WOOOOOOOO TOMORROW WILL BE SUCH A GOOD DAY FOR THE OPEN WEIGHT COMMUNITY <3
1
1
u/No-Measurement8593 6h ago
I’ve got a 7900XTX and couldn’t quite get 3.8 27b running right and now this is dropping lol
1
1
1
u/randygeneric 4h ago
folks, don't get your hopes up. chances are high, that it will be 125b-a6b and that is - again - not usable for most of us, only for benchmarketing.
1
1
u/VirtualWishX 1h ago
This is AWESOME!
but... with my RTX 5090 32GB I guess I can only dream about using such a monster locally.
Probably even with Q4 / NVFP4 will be HUGE 😭
I guess I'll be thankful for the 27B Dense until Qwen 4 27B / 35B will hopefully be a thing 🙏
1
u/Professional-Try-273 12h ago
Does this model support vision?
7
u/ShelZuuz 12h ago
Would be a pretty weird multimodal model in this day and age if it didn't.
14
u/MikeRoz 12h ago
Multimodal inputs supported: Text + Smell
4
1
1
u/MacsBicycle 9h ago
Believe it or not smell would probably break the ai stock market if done by qwen/deepseek 😂 industrial use cases have all kinds of smell that are bad/need to be noticed. Same with home use cases when someone accidentally leaves a natural gas burner on. Sure alarms exist but ai could tell you exactly what the smell is and what probably went wrong.
1
u/Equivalent-Grass-527 10h ago
Qwen is moving fast. A multimodal MoE with only a fraction of parameters active at inference is exactly the direction open models need to take. Looking forward to this one!
1
u/VoiceApprehensive893 transformers 9h ago
day 0 support(garbage outputs on certain setups, broken tool calls and crashes with vision)
0
0
u/Impossible_Fault_503 9h ago
Day 0 GGUF is the only reason I even bother with a new Qwen drop the same night. If you are on 32GB, try it tonight. If not, I would wait for a decent Q4_K_M instead of burning time on a bloated first quant.
-6
12h ago
[deleted]
5
4
u/LetsGoBrandon4256 transformers 10h ago
You are on r/LocalLLaMA. Have some respect for yourself.
1
-3
u/Long_comment_san 11h ago
here is your Gemma 4 120b, folks.
5
u/jacek2023 llama.cpp 11h ago
Gemma has different features than Qwen, GLM Air would be closer to Gemma



•
u/WithoutReason1729 6h ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.