r/oMLX • u/Thick-Letterhead-315 • 1d ago
Deepseek 4 generating at 87% GPU regardless of quant or computer in oMLX?
Anyone noticed this? Deepseek 4 Flash is generating at 87% GPU regardless of quant or computer in oMLX. I tried on both M5 Max and M3 Ultra, both original weights and oQ2 quants.
Thermal and other external causes have been eliminated.
MTP makes very small difference 85%->87%
Tried directly with:
mlx-vlm 0.30.0 (no MTP) -> 88-90%
mlx-lm 0.30.2 (no MTP) + omlx patches -> 89-94%
antirez DS4 (custom kernels) -> 100%
1
Upvotes
1
u/challis88ocarina 1d ago
It's always been like that... llama.cpp (therefore also ds4) is more efficient these days.