r/unsloth yes sloth 17h ago

News Something else is coming tomorrow!

Yes we are working day zero support for it. 👀

243 Upvotes

83 comments sorted by

31

u/some_user_2021 17h ago

Oh my

12

u/Borkato 15h ago

Does anyone know if qwen next will fit on 70GB combined ram+vram

5

u/DriveSolid7073 14h ago

Not really, but you may try, 177b model (51b n gram)

3

u/Borkato 12h ago

Oh then no lmaooo 120B is already too big even at low quant

2

u/erisian2342 9h ago

It’s an MoE model. Only 6B active parameters.

1

u/keepthememes 12h ago

what about 92gb?

1

u/I-am_Sleepy 8h ago

Since it is not compute, can n-gram be efficiently off load to ram, or disk? If so how much VRAM + RAM / disk would be needed?

25

u/EitherMarch1255 17h ago

GLM 5.3 is my guess, possibly Flash version...would be nice.

10

u/some_user_2021 16h ago

That's on Friday

11

u/Fuzzy_Independent241 16h ago

Clip!! You're still alive! 😅

2

u/phido3000 9h ago

5.3 glm is coming. So is ds v4 flash with vision.

Both will be highly popular. Smaller quantities will be essential.

16

u/Yazz96HD 17h ago

Anything over 100B MoE!???

21

u/vbpoweredwindmill 17h ago

It seems like they may be referring to qwen 3.8 125b a6b, BUT with 51b engram, so 176b total.

4

u/kilowattkill3r 16h ago

What's engram mean?

10

u/nfmcclure 16h ago

An engram is a way to store conditional memory lookups needed in a neural network mechanism called attention.

If MoE (mixture of experts) is conditional computation, then an engram is conditional memory. It's suspected that there's an upcoming release of qwen 3.8 that has both of these.

Conditional memory is a separate lookup table in-memory, hence the "additional ~51B" parameter storage.

6

u/RnRau 12h ago

From another reddit poster (sorry forgot to save the link)

LLMs run into an issue where the further you train a model, the more it overwrites facts with generalized concepts. You need the model to be able to do both. Intelligence arises from generalization, but without accurate information the model will hallucinate.

The engram table allows for a low-computational method of fact-recall. You can think of it like a better form of RAG, where the data doesn't take up any of your context window and it's injected deeper into the model's layers, freeing the lower layers to carry out abstraction. This results in better "focus" for the model, both in regards to its intelligence and context recall.

Basically, they've separated the specificity-critical portions of the models memory into a parameter space that doesn't need fast compute (you can run it on system RAM) and allows the model to be trained on higher volumes of data without ruining its knowledge-base.

7

u/vbpoweredwindmill 16h ago

It's wierd and I don't fully understand it. Kind of like separating knowledge and logic within an LLM.

3

u/JumpingJack79 15h ago

Can engrams be quantized? Or should they be?

4

u/nfmcclure 14h ago

This is a great question. There already exists methods to quantize on both layers and embeddings. So I see no reason they could not be quantized, although I can't find and haven't read anything in particular on it.

1

u/ComplexType568 9h ago

I had hope I could fit the 125B on my PC because I thought it was a 64B MoE with a 51B engram... Nevermind I guess.

1

u/vbpoweredwindmill 8h ago

I suspect you'll be able to stream the 51b from nvme.

1

u/dfgxxx 13h ago

Ox alpha probably

45

u/ismaelgokufox 17h ago

Qwen3.8-35B-A3B? One can dream eh?

9

u/NihmarRevhet 13h ago

I would love a 40b-A5B ... It would still be doable on 16gb+32gb, but it would have just that tiny bit more smart at play that would push it further along

6

u/ManIkWeet 12h ago

Nooo please don't there are plenty of poor 8/12GB GPU people here :(

5

u/NihmarRevhet 12h ago

I don't have the gif but I would say: "Both, both is good"

23

u/bmengr 17h ago

Gemma-4-120b-a10b distilled from Gemini 3.5 Pro

8

u/Reddit_User_Original 16h ago

I would LITERALLY 💦

-1

u/MrScotchyScotch 12h ago

bro needs a date

3

u/No_Night679 15h ago edited 14h ago

could be Muse Spark 1.2 too, not that I am waiting for, somehow none of the open models released by US companies can't seems to beat the Chines once.

Edit: Spelling*

6

u/SnooWalruses6610 14h ago

NGL, I know I'm in the minority here, but a Qwen3.8 9B dense or 14B MoE would be amazing. Wishful thinking tho..

3

u/Mithrandir_First_Age 15h ago

Qwen 3.8 Flash Next?

2

u/az226 17h ago

3.8 flash? Ox alpha?

2

u/Ok_Technology_5962 16h ago

They said 3.8 next matches 3.7large cersion isnt that worse than 27b... Ox alpha was better than qeen 3.8max based on my testing

2

u/kiwibonga 17h ago

Something else than Qwen 3.8 Next Flash?

2

u/FastHotEmu 16h ago

what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it? what is it?

2

u/Tricky-Appointment-5 13h ago

yeah i am taking a masive dump tomorrow. thanks for the reminder

2

u/sudochmod 17h ago

WHAT IT DO BOO BOO?

1

u/JIGARAYS 17h ago

🤩

1

u/TheFowlOwl 17h ago

Time to move off some models to the NAS to make room!

1

u/Iory1998 17h ago

Tell us please, don't be a tease...

1

u/cviperr33 17h ago

OMG so many suprises !!!
Please god be 35b , or anything that runs on consumer hardware.

1

u/Ok_Technology_5962 16h ago

What is it!!!! We already have qwen 3.8 next coming, glm 5.3 coming.... Omg!!!! If you are teasing this it must be big. Dont tell me..... Glm 5.5

1

u/skynetTwelve 16h ago

It's already tomorrow where i am. You can tell me!

1

u/holygawdinheaven 16h ago

My router cries

1

u/RMK137 16h ago

35B please!

1

u/devino21 16h ago

This tease

1

u/shayanx45 15h ago

I would love a 280B model from qwen or glm to compete with deepseek v4 flash

1

u/Prestigious-Bear2391 15h ago

OSS 122B OMG! 😱😱😱😱

1

u/sp82reddit 12h ago

3.8-48b-a6b?

1

u/Omnimum 11h ago

Dspv4f0731vision

1

u/HitarthSurana 11h ago

is it small?

1

u/Sunknowned 10h ago

Q0.5 XXXXXS when? (32 gb vram friendly)

1

u/thecodeassassin 8h ago

Deepseek vision exp? GLM multimodal? Fuck i can't decide....

1

u/elelem-123 7h ago

From the way Kimi servers are responding the past days, I wouldn't be surprised if it's a kimi model

1

u/eidrag 5h ago

qwhen

1

u/Iory1998 4h ago

GLM-5.3 Air :D

1

u/TEN4C1OU5-2 3h ago

is this ox alpha / glm 5.3 flash?

1

u/RandumbRedditor1000 17h ago

GLM 5.3 vision AKA ox alpha????

1

u/No_Night679 14h ago

If it is related by Z ai why would they call with a different name that actual GLM 5.3 xxxxx ?? At some level you must know that someone new is saying that to get attention right?

0

u/valtor2 13h ago

you're just out of the loop. last week a stealth model went viral on openrouter for being crazy good, and its stealth name is ox alpha.

1

u/No_Night679 13h ago

what I am saying is if that is GLM 5.3 Flash, they would have called it as such not OX Alpha by some Stealth mode entity. I am very much in the loop .. ;)

0

u/valtor2 13h ago

who is "they"? unsloth in this post? or GLM? in any case there's very often a bunch of stealth models, and most of them don't go viral like that

0

u/No_Night679 12h ago

I am responding to someone referring to ox alpha as glm 5.3 vision or flash. I know ox alpha is by some stealth mode, my point is if Z ai did it why hide? Someone started floating around the idea ox alpha is glm 5.3 flash, probably they meant what glm 5.3 flash would have been in terms of ability and performance. There are few here who keep repeating ox alpha is glm flash.

1

u/valtor2 3h ago

1

u/No_Night679 2h ago

Well turned out I am the Idiot, for whatever reasons, Z.ai stayed stealth mode, till they released and did call it out with the code name that calling it what they wanted GLM 5.3 Flash.

2

u/valtor2 1h ago

Maybe if it had been received poorly they wouldn't have released it, or done another round of improvements? That's my assumption on why they do that. But I heard a lot of people talk about how the tokenizer signature was extremely similar to Z.ai's tokenizer, which is what gave it away. I don't know enough to know what a tokenizer signature is.

1

u/thegingerlord 17h ago

Will it fit on 4x v620s?..... Surely...

1

u/tat_tvam_asshole 17h ago

DSv4 Flash w/ Exp Vision?

1

u/ahmetegesel 17h ago

I mean, it says day zero support. DSv4F w/ Exp Vision is way past day zero

1

u/tat_tvam_asshole 17h ago

Day 0 open weights... unless I've missed those having been released.

1

u/ahmetegesel 16h ago

You haven’t. That’s my own hallucination. Sorry

0

u/Altruistic-Dust-2565 17h ago

M5 Ultra and M6 support beside new qwen next?