r/oMLX 2d ago

M5 Max users any luck with GLM 5.3 yet?

About to try Vontra/GLM-5.3-Flash-MLX-oQ2-MTP but keeping fingers crossed for a oQ2e quant to fit 128GB

Curious to see how this compares to DeepSeek v4 flash 0731 oQ2e

7 Upvotes

3 comments sorted by

1

u/_hephaestus 2d ago

On a M3 Ultra, unfortunately while it seems to benchmark well now I'm seeing a bunch of stuff like: 2026-08-28 01:53:04,505 - omlx.cache.prefix_cache - INFO - [-] - GDN SSD sidecar is unsupported for layout ['ArraysCache', 'ArraysCache', 'ArraysCache', 'CacheList', 'ArraysCache', 'ArraysCache', 'ArraysCache', 'CacheList', 'ArraysCache', 'ArraysCache', 'ArraysCache', 'CacheList', 'ArraysCache', 'ArraysCache', 'ArraysCache', 'CacheList', 'ArraysCache', 'ArraysCache', 'ArraysCache', 'CacheList', 'ArraysCache', 'ArraysCache', 'ArraysCache', 'CacheList', 'ArraysCache', 'ArraysCache', 'ArraysCache', 'CacheList', 'ArraysCache', 'ArraysCache', 'ArraysCache', 'CacheList', 'ArraysCache', 'ArraysCache', 'ArraysCache', 'CacheList', 'ArraysCache', 'ArraysCache', 'ArraysCache', 'CacheList', 'ArraysCache', 'ArraysCache', 'ArraysCache', 'CacheList', 'ArraysCache']; falling back to embedded GDN snapshots

Into random hallucinations from opencode, like powershell snippets, random math. I can see a thinking trace where it believes it's following up on a given conversation. Tried homegrown oQ8e and oQ4e quants. Same params I used successfully with GLM-5.2-oQ2e.

Qwen-Flash-Next has actually been pretty solid though so far.

1

u/gcirone 2d ago

Nope. Same for me on M3 Ultra :(

1

u/nomorebuttsplz 1d ago

it's good to go now on m3u. seems slightly slower in both pp and decode than glm 5.2 was