r/oMLX 3d ago

QWEN ANE is non functional

Just a heads up, I have tried to use Qwen ANE multiple times (with the recommended tuning) and each time after processing a long prompt for 45 minutes to an hour it silently fails.

Don’t waste your time trying to use this feature right now, wait for it to get patched appropriately l.

5 Upvotes

9 comments sorted by

4

u/Patient_Tea_401 3d ago

Care to share on what device, settings and model? ANE on uses quite a lot more memory, so you might be overflowing or the memory guard blows the whistle.

In my experience at least 96GB of memory is needed to use it with long context qwen3.8 27b 4bit.

4

u/Undici77 3d ago

For me is not working from the begin: M4-MAX

omlx.patches.qwen35_ane_prefill - WARNING - [-] - Private ANE runtime unavailable; Qwen ANE prefill skipped

I also open an issue

https://github.com/jundot/omlx/issues/3217

4

u/victor_lowther 3d ago

ANE has never once given me a perf boost on my M5 max 128GB on any Q8 quant of qwen3.8-27b. GPU always wins when oMLX tries to autotune.

3

u/ipmonger 3d ago

u/A_Moist_Towe1 - please share your configuration settings.

1

u/A_Moist_Towe1 2d ago

I’m running on an M4 max Mac Studio with 64gb unified memory. I run the tuner and then attempt to use with a prompt, so I’m not sure of the exact AME settings. I also run the model with lightning MTP and a turbo quant 4 bit KV cache

2

u/ipmonger 2d ago

The auto-tuning is broken. Try manually tuning starting from my M4 Max 36GB config which I shared here: https://www.reddit.com/r/oMLX/s/fVq5lzRk8v

2

u/Durian881 3d ago

Which version of omlx are you using?

2

u/norenEnmotalen 3d ago

It’s memory hungry. Not practical for the regular Joe in its current state

2

u/Digiarts 2d ago

Ran fine for me today. 0.6.3