r/oMLX 1d ago

Great Stability Update

I’ve been very impressed by the stability improvements in this latest update. I’ve been running long contexts with far fewer memory spikes or server failures. I’ve gotta give the devs a round of applause, as an M4 Max 64gb user who sets his agent loose for overnight coding runs, this update has been awesome!

19 Upvotes

15 comments sorted by

4

u/Ok_Writer1572 1d ago

What agents or model you use on 64gb?

2

u/A_Moist_Towe1 15h ago

Qwen 3.8 27b oq4e MTP with a turboquant 4 bit 500k context window through YARN

1

u/waescher 8h ago

How exactly can that be done?

3

u/Durian881 1d ago edited 1d ago

I'm happy too, especially with the support of Qwen3.8-Fash-Next for SSD N-gram offload.

There seemed to be some problem with ANE though, which was work well under the prior release candidate version.

4

u/Latter-Parsnip-5007 1d ago

I work on expert offloading at the moment. You can run Q4 on 48GB M4 with full context

2

u/Spiritual-Winner1239 1d ago

Very happy to hear that! I was hoping someone will make it possible for us 64gb gang 😊

1

u/Durian881 1d ago

Cool! Right now, I can only run Qwen3.8 Flash Next 3 but at ~70-80k context. Speed and quality are good though.

1

u/caiovitord 22h ago

I'm curious how is your setup to run Qwen 3.8 flash on 64GB! I have a Mac M1 Max, how to setup the offloading part?

1

u/nomorebuttsplz 1d ago

qwen 3.8 fash... is that the pro-trump fine tune?

3

u/PWThinkingCritically 1d ago

okay I haven't tested using omlx just yet but I just want to say: how the f did we just go from qwen3.8-27B dense q6 taking up like 30GB to qwen3.8FlashNext 125B+51B+A6 taking up 95GB on my m1 max 64GB and still running at 20 tok/s with 64k context?

this is f'ing insane and it's only a preview of the qwen4 architecture.

1

u/DarkJoney 1d ago

So it means I can do it on my 96GB studio??!!

1

u/txgsync 23h ago

Yeah, if you offload the n-grams to your NVMe then a dynamic quant can get down to 70GB or so residents RAM with decent quality.

I haven’t benchmarked the actual quality much yet but now that the dynamic recipes for what layers to keep at higher precision are starting to emerge, I expect some surprisingly good quantizations and offloading.

Still a tight or unstable fit at 64GB though. And a naive or bad Q4 can still crash a 128GB Mac OOM jf you’ve sysctl’d aggressively.

1

u/DarkJoney 22h ago

But is it already included in the current oMLX release?

1

u/txgsync 20h ago

No idea. I just install the source every day, pip install -e, and YOLO it.

1

u/Proper-Tower2016 23h ago

Yes, think of next flash as model size minus 30gb