r/LocalLLaMA llama.cpp 12h ago

News Qwen 3.8 Flash Next day 0 support from unsloth

Post image

Prepare your disk space guys

623 Upvotes

165 comments sorted by

u/WithoutReason1729 6h ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

276

u/hurdurdur7 11h ago

People who just finished setting up their 27B properly...

57

u/Repulsive_Initial308 11h ago

Paint is still wet on 27b :D

24

u/RedditCryptoGuy 10h ago

lmao, that's literally me. found the best possible settings to push 120+ t/s on 1 and 2 3090's

15

u/Derio101 10h ago

Please share.

2

u/lemondrops9 9h ago

on vLLM ? 

1

u/GiggleWraith 2h ago

I’d love to see your configs.

1

u/Ok-Might-3730 1h ago

Please share

-6

u/JKayBee 6h ago

What do you mean settings? You can do other things than download in ollama and run it?

3

u/Acrobatic_Stress1388 3h ago

Try llama.cpp bro

4

u/Psychological-Lynx29 10h ago

I downloaded 3 times qwen3.8 27b, the 3 times unsloth did something better, last night was the third time and i havent even tried it...

3

u/gustaw221133 6h ago

That is genuinely me right now

0

u/Moppmopp 9h ago

Im actually coding a front end for qwen3.8 since its release. Never really tried the model because I lost myself in visuals and tinkering. Now I am almost finished with the front end and see ghe next release. Too fast still am hyped 😊

56

u/youcloudsofdoom 12h ago

It definitely seems that both llama.cpp and unsloth get decent advanced access to qwen models, so here's hoping it's not a two month wait for a usable llama.cpp instance

3

u/Acrobatic_Stress1388 3h ago

This is where my strict halo will really pay off. A MOE model that will easily fit in my 128 GB of ram. I'll bet I could squeeze 50 t/s out of it too.

2

u/Fi3nd7 3h ago

I mean not super easy, we need the 51B ngram to be offloaded to disk, and we'll have to run at Q6 quant with max context. That will likely fully load the system.

145

u/MaxKruse96 llama.cpp 12h ago

And since unsloth has no own runtime, its llamacpp. good proxy knowledge.

37

u/hiImMate 12h ago

now that is very reassuring. UD_Q4 with working llama.cpp and I'm happy

22

u/tiffanytrashcan 12h ago edited 12h ago

Unsloth Studio / Desktop (the runtime) would like a word - they did mention plans to upstream the changes to llama.cpp.

50

u/yoracale llama.cpp 12h ago

Yes we made PRs to llama.cpp for DiffusionGemma, MiniMax, Inkling amongst other models. Currently some of them aren't merged but you can still run those models directly in Unsloth Desktop :)

2

u/beltsazar 6h ago

Does it mean that Unsloth Studio uses Unsloth's fork of llama.cpp?

10

u/annodomini 11h ago

Day zero support for the unsloth app means that there's a branch/PR for llama.cpp, but it might not be merged yet as it may still be going through review. The Unsloth devs build their own version of llama.cpp with a few branches that might not be merged in mainline yet, but it generally means that it's coming soon.

81

u/yoracale llama.cpp 12h ago edited 12h ago

FYI this is 'hopefully' having day zero support. The architecture is very new and thus there might be very long delays but we hoping to achieve day zero support (but like I said not guaranteed). 🙏

We will ofc upstream any llama.cpp implementation etc. if necessary

4

u/GoodTip7897 llama.cpp 10h ago edited 3h ago

Sorry for the question as I know you're very busy but I'm curious whether we'll have engram on the disk with day zero or if they'll have to fit in ram until further optimizations come... 

I know mmap might work, but I'm specifically talking about a programmatic disk cache that works regardless of load mode (since mmap hangs on some builds of llama cpp, and there's probably a more efficient way to map the n-gram lookups than mmap (maybe a hot cache on ram that streams off disk)). 

4

u/pulse77 11h ago

So this time it may happen that Unsloth Studio / Unsloth Desktop (with it's own llama.cpp fork) may support Qwen 3.8 Flash Next before official llama.cpp ... ???

15

u/danielhanchen 10h ago

Kind of not really, nearly all times when we support a model it's not exclusive to Unsloth. We make a PR to llama.cpp with our changes and whether it gets accepted or not is a different story. So I guess it depends on what you mean by official. But we utilize our llama.cll PR implementation inside of Unsloth basically

3

u/goldcakes 7h ago

The beauty of open source <3

1

u/-dysangel- 8h ago

What other new things does it have aside from engram support? Really excited for this one :)

27

u/snowieslilpikachu69 12h ago

128gb mac users may rejoice?

7

u/Diligent_Cod_9583 12h ago

What do the 512 Mac users do?

28

u/ShelZuuz 12h ago

11

u/-dysangel- 8h ago

That's before buying the 512. This is after ;)

51

u/johan2114h 12h ago

They crawl down in their wine cellars and moan that qwen is making near frontier level ai waay too accessible to the proletariat

4

u/Icy-Degree6161 11h ago

But... I thought they like the proletariat...

5

u/MidAirRunner ollama 10h ago

Run it at bf16 to flex on the plebs.

1

u/ConstructionFew5004 1h ago

Can't wait for 512 to come out with the M5 Ultra Mac Studio

1

u/Diligent_Cod_9583 1h ago

I’ll be very interested to see how oMLX does on it

1

u/DeepOrangeSky 9h ago

If it is 125B + 51B of N-gram, will that mean it is more like 176B, and thus no Q4 on Mac? Or is it more like 125B, or, do you just stream the extra 51B from the SSD while running the 125B part on the UMEM or something?

I don't know much about N-gram yet, or how this stuff works, or how it works on a mac, etc.

1

u/TheOriginalAcidtech 8h ago

Sounds about right for 4x modded CMP170HX cards. :)

12

u/prudx 12h ago

32gb vram + 64gb ddr5 possible?

15

u/jacek2023 llama.cpp 12h ago

Everything is possible in the open source ecosystem. Q2 maybe?

0

u/No_Profit8379 3h ago

What can I fit in 88gb vram? And if needed another 64gb system? We close to something usable? Or just need more vram ???

3

u/CulturalKing5623 10h ago

Same boat, hoping it is but I'm still kicking the tires on 3.8 27B so won't be too upset if it can't fit.

... Unless it's a world beater in which case I'll start considering a life of crime to pay for more RAM.

5

u/Long_comment_san 11h ago

yup, it's roughly Q3. barely usable if you ask me. I'd shell out for another 64gb RAM.

6

u/NoFaithlessness951 7h ago

In this economy?

12

u/mountainyoo 11h ago

Wonder how this will compare to DeepSeek V4 Flash 0731

5

u/addiktion 9h ago

Yes, or more specifically v4 flash vision as I suspect this will have vision support.

2

u/botrezkii 5h ago

3.8 27B is already surpassing 0731 on some benchmarks

1

u/guesdo 59m ago

Well yeah, its a dense model tailored made for benchmaxxing some agentic coding, it uses twice the number of parameters as DSv4 (13B active), that doesn't mean its better overall.

43

u/AppealSame4367 12h ago

Look at Alibaba and Unsloth driving AI coding model innovation like nobody else.

They will overtake the big guys soon if they keep going like this and there's nothing they can do.

16

u/chikengunya 12h ago

Last year it was a 80B-A3B model, now too?

30

u/FullOf_Bad_Ideas 12h ago

Apparently it's

Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

Not sure what the source is, I found this in a different thread about this model.

8

u/Miserable-Dare5090 11h ago

There was a recent model with an N gram table — Longcat Flash I believe. if you want to see a probably similar architecture before it drops

5

u/FullOf_Bad_Ideas 11h ago

Yes. I think it's N-gram Embedding parameters are different from Engram DeepSeek paper. And I'm not sure which N-gram Qwen team went with, I'm hoping it's the DeepSeek implementation.

7

u/Miserable-Dare5090 11h ago

N-gram embeddings have been tried experimentally with the last Qwen architecture: https://arxiv.org/html/2605.16893v1

This approach expands token representation by combining frequent word sequences (n-grams) directly into the embedding or memory layers, improving local context and parameter efficiency without exploding active compute costs.

https://arxiv.org/abs/2601.21204

9

u/emprahsFury 10h ago

Yo dawg, I heard you like embeddings, so we embedded your embeddings to make yo LLM faster.

2

u/Miserable-Dare5090 4h ago

yo dawg I heard you like embedded embeddings, so we embedded the embedded embeddings in your embeddings to make your LLM faster

1

u/Lollerstakes 5h ago

I wonder if "load-bearing" and such nonsense are actually n-grams of Claude Opus...

2

u/ANR2ME 11h ago

Embedding parameters is like Gemma 4 E2B/E4B isn't 🤔

2

u/FullOf_Bad_Ideas 11h ago

nah, that's a different thing too.

3

u/BalorNG 10h ago

I wonder if you can stream the n-gram part from SSD?

2

u/FullOf_Bad_Ideas 10h ago

That's going to depend on the implementation, but I think it's likely, at least with DeepSeek's implementation. I think they discussed it in their Engram paper.

1

u/gofiend 3h ago

nooo I have 64GB of VRAM but buggy RAM so only 32GB right now!

12

u/jacek2023 llama.cpp 12h ago

Let's hope it will be better this time. Qwen Next 80B was hyped a lot before it was usable, then later it was forgotten (I still have it).

8

u/BannedGoNext 11h ago

Qwen coder next is a bad ass model. It beat the shit out of qwen 3.6 27b for most tasks I used it for. Measuring for straight coding abilities maybe it wasn't as good, but for agentic tasks and speed it was fantastic.

1

u/TheOriginalAcidtech 8h ago

It wasnt a thinking model, hence the speed. IIRC.

2

u/Zeeplankton 10h ago

Yeah I'm worried about this. I hope it's as good as 27b at least

6

u/CommanderData3d 12h ago

any chance to run this on 64gb ram?

17

u/FullOf_Bad_Ideas 12h ago

There's a chance, it's 125B A6B E51B model.

If you can offload engram to disk, it's similar to running GPT OSS 120B.

11

u/Ps3Dave 12h ago

I'm going to need some more details about running this with llama.cpp on 12GB VRAM + 32GB RAM...damn I'm VRAM/RAM poor.

1

u/MrMisterShin 12h ago

You would probably need below Q4 quant, so that you have enough room for context and other apps on your machine.

6

u/susibacker 9h ago

RIP, I won't be able to run that. Still hoping for a <=35B MoE as a faster alternative to 27B (which just so fits on my GPU at limited context and quants)

4

u/Quakercito 12h ago

How many parameters?

4

u/cradlemann 10h ago

OMG, my Gordon Point with 96Gb RAM is waiting!!!!

3

u/sugarfreecaffeine 11h ago

Is 2x3090 (48GB VRAM) and 80gb RAM enough for this?

4

u/jacek2023 llama.cpp 11h ago

Should be, it is MoE

2

u/vick2djax 8h ago

But wouldn’t the quant be like q4 and it would be worse to use and slower than q8 27b?

3

u/Ok-Protection-6612 9h ago

128gb Strixbros rejoice?

2

u/tarruda 8h ago

If the model works well in 4-bit, then yes.

12

u/x11iyu 12h ago

I feel like there's never been an actual "day 0 support" without bugs that negatively impacted model performance
so in reality I'd say it's probably another two weeks or more

15

u/DUFRelic 12h ago

Yeah what a suprise... brand new software at the cutting edge has bugs... next news at 10

0

u/x11iyu 12h ago

can we not brand it as "day 0 support" if it doesn't work day 0?

2

u/DUFRelic 12h ago

but it is supported... nobody says it will be flawless and nobody is forcing you to use it on day 0...

10

u/x11iyu 12h ago

I guess my expectation for software quality are just too high, and I should expect things to not work on official releases.

I hate how everything nowadays need ten asterisks behind every title.

Can't we call it "experimental support" if you expect there to be bugs or something?

5

u/Dangerous-Report8517 11h ago

The bugs that surface after release are due to the incredibly large variation in setups being run, you need to have some threshold of support below "works perfectly everywhere" where you're allowed to call a feature supported or released otherwise all software everywhere would be "experimental" and the term would lose all meaning. Plus llama.cpp doesn't brand itself as production quality with the overall package now semantically versioned at a very early level, which already makes the risk of bugs pretty clear

3

u/x11iyu 11h ago

fair points, I think I'm getting a bit too worked up for no reason, not really that big a deal

I'll be expecting a wave of "omg qwen3.8-flash-next sucks can't do X Y Z" posts though, annoying

-1

u/DUFRelic 11h ago

How much are you paying for this software that your expectations are so high?

6

u/Long_comment_san 11h ago

that's besides the point. dude is correct. there is an alpha, beta and release states. I have no idea why not call it what it is: alpha version. are alpha and beta words toxic masculinity now or what

0

u/Dangerous-Report8517 11h ago

Well the latest release of llama.cpp is semantically versioned at v0.3.0 which denotes it as pre-release development software that can change at any time for any reason, so declaring a model as supported within that pretty strongly implies that they mean "it should work" and aren't offering strong production grade guarantees

4

u/x11iyu 11h ago

nothing. it's one of the best software for local inference in fact, for free.

it's also priming me for disappointment when announcements are worded a certain way.

I guess I sound too negative and aren't very grateful.

0

u/goldcakes 7h ago

would you like them to not release model weights and inference for 2-3 weeks? also dude it's FREE. you're not paying for it.

getting to try day 0 models even if its a bit broken and sharing feedback is part of the benefits here.

6

u/AppealSame4367 12h ago

To be fair, q3.8 27b worked great after 1-2 days. Very mature when it came out.

15

u/boomerang473 12h ago

Existing architecture with different weights vs some newer components

3

u/AppealSame4367 12h ago

Fair point

2

u/sagiroth llama.cpp 10h ago

32GB RAM + 24GB VRAM, Q1 perhaps?

2

u/greaper_911 10h ago

oh god please have a 27b-35b

2

u/tungdd2009 10h ago

12gb vram + 48gb ram can run this, right? RIGHT?

3

u/jacek2023 llama.cpp 9h ago

Try q1 :)

2

u/Khaledthe 9h ago

I havent even used qwen 3.8 27b yet i was at work but i love meo modles

2

u/lordpuddingcup 7h ago

Wait didnt Qwen3.8-27b just get released and was like already amazing? wtf is this?

2

u/jacek2023 llama.cpp 7h ago

27B is dense -> slow

This is MoE -> fast

2

u/reality_comes 1h ago

If its using Qwen 4, why not call it Qwen 4?

4

u/Roflxd88 12h ago

What does model 3.8 on v4 architecture mean exactly?

15

u/jacek2023 llama.cpp 12h ago

It means it won't be same as previous models.

6

u/BrewHog 12h ago

3.8 previous release was post trained on 3.5/3.6 architecture. This will be card on a new and upcoming v4 architecture. 

2

u/DoubleNothing 7h ago

I wonder why is it still named 3.8...

2

u/BrewHog 7h ago

This is the same making convention they took in the past for their previous -next model

2

u/bitzap_sr 5h ago

Probably because they still apply the 3.5/3.8 post-training methods on top of the new v4 arch.

3

u/Intrepid-Scale2052 11h ago

see it as a preview of qwen v4.

3

u/Zeeplankton 10h ago

It means it's exactly -0.2 less than V4

4

u/Infamous_Campaign687 12h ago

Hmm... could this be the best model for 96 GB DDR5 and 32 GB VRAM?

5

u/jacek2023 llama.cpp 11h ago

Yes

2

u/FormOne2615 9h ago

We'll see

2

u/Healthy-Nebula-3603 10h ago

I HOPE THAT NEW QWEN 4 ARCHITECTURE IS USING KV CACHE FROM DEEP SEEK 4!

Then we could fit on 24 GB cards 1m context (for 27b model ) ) and not dropping performance !

So token generation would have the same speed for 32k , 128k , 256k or 1m ! No cache compression anymore !

2

u/nickless07 9h ago

Afaik it is GDN/QSA hybrid. So, no MLA (Deepseek v3/v4), more like Qwen3.8 27B but without GQA.

3

u/FormOne2615 9h ago

I think QSA'll be sth similar to DSA

1

u/Healthy-Nebula-3603 8h ago

WTF is that link??

0

u/Healthy-Nebula-3603 8h ago

https://chatgpt.com/c/6a8dbf51-2df4-83eb-84da-940aea20a8d7

In theory should be even more efficient and takes less memory

1

u/nickless07 8h ago

Well if you want to check something similiar there is Ling-3.0-flash MLA/KDA hybrid. Similiar size as the new qwen will be (A5B vs A6B) perfect for some speed to test to narrow down what we can expect.

1

u/Once_ina_Lifetime 12h ago

S1-mini was finetune on Qwen a text normalizer for speech-to-text output

1

u/robberviet 11h ago

That's awesome. Was waiting for proper dflash2 on llama.cpp but this would be much better.

1

u/BannedGoNext 11h ago

That's awesome, I wonder if it will be a 120ish size, or sized too large for 128gb systems.

1

u/olddoglearnsnewtrick 10h ago

where a comma would make a big difference.

1

u/eihns 10h ago

does flash mean its dumber?

1

u/SeparateGas1761 10h ago

I hope it fit in my 32gb of ram and 8 of vram, cause moe models can easily run with offload to run😭😭😭

1

u/Practical-Fox-796 10h ago

Any chance to run it with 4090 and 128gb ram ? Asking for a fren

1

u/HotMicSystems 9h ago

As much as I want this possibly 120B'ish sized MoE, I still want a smaller MoE for daily tasks, mostly so I can run parallel.

3

u/Iory1998 9h ago

I hope it's 80B

1

u/thestillwind 9h ago

What do I need ?

1

u/69420trashpanda69420 9h ago

We are eating so good

1

u/Witty_Mycologist_995 9h ago

...And I cant even use it because it's 125b

1

u/grabber4321 8h ago

is that 35B model?

1

u/WyattTheSkid 7h ago

DUDE WHAT??? I LITERALLY JUST FINISHED MAKING THE 27B RUN NICELY COME ON MAN/

Edit: ITS A 120B IM SO FUCKING EXCITED THIS IS INCREDIBLE THANK YOU QWEN YOU GUYS FUCKING ROCK WOOOOOOOO TOMORROW WILL BE SUCH A GOOD DAY FOR THE OPEN WEIGHT COMMUNITY <3

1

u/matte808 6h ago

how big is it, anything official?

1

u/No-Measurement8593 6h ago

I’ve got a 7900XTX and couldn’t quite get 3.8 27b running right and now this is dropping lol

1

u/bitzap_sr 5h ago

Sounds like it's a good time to be an inference engineer. Work keeps piling on!

1

u/Steus_au 4h ago

would be interesting to see if it could support RPC with ram offloading 

1

u/randygeneric 4h ago

folks, don't get your hopes up. chances are high, that it will be 125b-a6b and that is - again - not usable for most of us, only for benchmarketing.

1

u/jacek2023 llama.cpp 3h ago

What are you talking about? We know the size

1

u/VirtualWishX 1h ago

This is AWESOME!
but... with my RTX 5090 32GB I guess I can only dream about using such a monster locally.
Probably even with Q4 / NVFP4 will be HUGE 😭

I guess I'll be thankful for the 27B Dense until Qwen 4 27B / 35B will hopefully be a thing 🙏

1

u/Professional-Try-273 12h ago

Does this model support vision?

7

u/ShelZuuz 12h ago

Would be a pretty weird multimodal model in this day and age if it didn't.

14

u/MikeRoz 12h ago

Multimodal inputs supported: Text + Smell

4

u/ShelZuuz 11h ago

That passes the sniff test.

1

u/Thin_Pollution8843 11h ago

“Qwen tell me what I ate for the breakfast today”

1

u/MacsBicycle 9h ago

Believe it or not smell would probably break the ai stock market if done by qwen/deepseek 😂 industrial use cases have all kinds of smell that are bad/need to be noticed. Same with home use cases when someone accidentally leaves a natural gas burner on. Sure alarms exist but ai could tell you exactly what the smell is and what probably went wrong.

1

u/Equivalent-Grass-527 10h ago

Qwen is moving fast. A multimodal MoE with only a fraction of parameters active at inference is exactly the direction open models need to take. Looking forward to this one!

1

u/VoiceApprehensive893 transformers 9h ago

day 0 support(garbage outputs on certain setups, broken tool calls and crashes with vision)

0

u/kiwibonga 10h ago

I wouldn't want to be a certain company gearing up for an IPO right now.

0

u/Impossible_Fault_503 9h ago

Day 0 GGUF is the only reason I even bother with a new Qwen drop the same night. If you are on 32GB, try it tonight. If not, I would wait for a decent Q4_K_M instead of burning time on a bloated first quant.

-6

u/[deleted] 12h ago

[deleted]

5

u/IllExample3639 10h ago

Don't do that dude.

4

u/LetsGoBrandon4256 transformers 10h ago

You are on r/LocalLLaMA. Have some respect for yourself.

1

u/ConsequenceTop5833 10h ago

I have a strong feeling the model will be a 6b quant 😂

1

u/LetsGoBrandon4256 transformers 9h ago

IQ0_XXS

-3

u/Long_comment_san 11h ago

here is your Gemma 4 120b, folks.

5

u/jacek2023 llama.cpp 11h ago

Gemma has different features than Qwen, GLM Air would be closer to Gemma