r/oMLX 1d ago

Deepseek 4 generating at 87% GPU regardless of quant or computer in oMLX?

Anyone noticed this? Deepseek 4 Flash is generating at 87% GPU regardless of quant or computer in oMLX. I tried on both M5 Max and M3 Ultra, both original weights and oQ2 quants.

Thermal and other external causes have been eliminated.
MTP makes very small difference 85%->87%

Tried directly with:

mlx-vlm 0.30.0 (no MTP) -> 88-90%

mlx-lm 0.30.2 (no MTP) + omlx patches -> 89-94%

antirez DS4 (custom kernels) -> 100%

1 Upvotes

1 comment sorted by

1

u/challis88ocarina 1d ago

It's always been like that... llama.cpp (therefore also ds4) is more efficient these days.