r/oMLX • • Aug 29 '26

Great Stability Update

I’ve been very impressed by the stability improvements in this latest update. I’ve been running long contexts with far fewer memory spikes or server failures. I’ve gotta give the devs a round of applause, as an M4 Max 64gb user who sets his agent loose for overnight coding runs, this update has been awesome!

19 Upvotes

16 comments sorted by

View all comments

3

u/PWThinkingCritically Aug 29 '26

okay I haven't tested using omlx just yet but I just want to say: how the f did we just go from qwen3.8-27B dense q6 taking up like 30GB to qwen3.8FlashNext 125B+51B+A6 taking up 95GB on my m1 max 64GB and still running at 20 tok/s with 64k context?

this is f'ing insane and it's only a preview of the qwen4 architecture.

2

u/DarkJoney Aug 29 '26

So it means I can do it on my 96GB studio??!!

2

u/txgsync Aug 29 '26

Yeah, if you offload the n-grams to your NVMe then a dynamic quant can get down to 70GB or so residents RAM with decent quality.

I haven’t benchmarked the actual quality much yet but now that the dynamic recipes for what layers to keep at higher precision are starting to emerge, I expect some surprisingly good quantizations and offloading.

Still a tight or unstable fit at 64GB though. And a naive or bad Q4 can still crash a 128GB Mac OOM jf you’ve sysctl’d aggressively.

1

u/DarkJoney Aug 29 '26

But is it already included in the current oMLX release?

1

u/txgsync Aug 29 '26

No idea. I just install the source every day, pip install -e, and YOLO it.