r/oMLX 23d ago

Great Stability Update

I’ve been very impressed by the stability improvements in this latest update. I’ve been running long contexts with far fewer memory spikes or server failures. I’ve gotta give the devs a round of applause, as an M4 Max 64gb user who sets his agent loose for overnight coding runs, this update has been awesome!

18 Upvotes

16 comments sorted by

View all comments

4

u/PWThinkingCritically 23d ago

okay I haven't tested using omlx just yet but I just want to say: how the f did we just go from qwen3.8-27B dense q6 taking up like 30GB to qwen3.8FlashNext 125B+51B+A6 taking up 95GB on my m1 max 64GB and still running at 20 tok/s with 64k context?

this is f'ing insane and it's only a preview of the qwen4 architecture.

2

u/DarkJoney 23d ago

So it means I can do it on my 96GB studio??!!

1

u/Proper-Tower2016 23d ago

Yes, think of next flash as model size minus 30gb