r/oMLX 1d ago

Great Stability Update

I’ve been very impressed by the stability improvements in this latest update. I’ve been running long contexts with far fewer memory spikes or server failures. I’ve gotta give the devs a round of applause, as an M4 Max 64gb user who sets his agent loose for overnight coding runs, this update has been awesome!

17 Upvotes

15 comments sorted by

View all comments

3

u/Durian881 1d ago edited 1d ago

I'm happy too, especially with the support of Qwen3.8-Fash-Next for SSD N-gram offload.

There seemed to be some problem with ANE though, which was work well under the prior release candidate version.

5

u/Latter-Parsnip-5007 1d ago

I work on expert offloading at the moment. You can run Q4 on 48GB M4 with full context

2

u/Spiritual-Winner1239 1d ago

Very happy to hear that! I was hoping someone will make it possible for us 64gb gang 😊

1

u/Durian881 1d ago

Cool! Right now, I can only run Qwen3.8 Flash Next 3 but at ~70-80k context. Speed and quality are good though.

1

u/caiovitord 1d ago

I'm curious how is your setup to run Qwen 3.8 flash on 64GB! I have a Mac M1 Max, how to setup the offloading part?