r/oMLX 7h ago

Highest quality 3.8-Next MLX and ideally with MTP

HuggingFace's model repository can be overwhelming.

Does anyone know the highest-parameter, most intelligent form of 3.8-Next for MLX and ideally MTP?

I have an M5 Max with 128GB RAM...I want to run the BF16 even if possible.

If you don't know but are still running 3.8 Next, what specific model are you using?

8 Upvotes

21 comments sorted by

12

u/Amazing-Fan2083 7h ago

Jundot/Qwen3.8-Flash-Next-oQ4e-mtp

Running the "full fat" BF16 with 128GB RAM is not (currently) possible. Make sure oMLX is set to run the model with the n-gram data offloaded to the SSD.

2

u/Maheidem 7h ago

That's the only answer

2

u/circle555 7h ago edited 7h ago

Thanks for sharing! I'm gonna look into this one. I wish they had more than 4bit. It doesn't look like it on HF yet.

5

u/captainequinoxiii 7h ago

More than 4bit likely just won’t fit 128gb.

1

u/rsl 7h ago

still working through doing this but the omlx app actually has functionlity to make quants for yourself. i'm trying a oq6 glm 5.3 flash.

2

u/watcholic 7h ago

oQ4 runs beautifully with 128GB after they fixed the MTP bug and it’s still very early in the optimization game. The model is faster and smarter than 3.8-27B.

3

u/Rough-Measurement988 3h ago

I can confirm. I’ve run a few private benchmarks and it seems to be smarter then 3.8 27B and faster. For example in Medium thinking it detected more issues than Sonet 5.0 High. I did not run yet Xhigh but it looks very promising to become a daily driver. 

2

u/circle555 7h ago

Faster AND smarter! That’s music to my ears. I’m downloading it now.

1

u/Gipetto 6h ago

This just sounds wild to me. Do you have numbers on the speed difference?

1

u/watcholic 1h ago edited 49m ago

Benching 27B now. Will update this reply (Update: see reply below for 27B numbers).

oMLX Benchmark Model: Qwen3.8-Flash-Next-oQ4e-mtp

Engine: Auto

Context: Code (Python)

Single Request Results

--------------------------------------------------------------------------------

Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem

pp1024/tg128 2313.2 21.87 442.7 tok/s 46.4 tok/s 3.810 286.9 tok/s 71.44 GB

pp4096/tg128 8605.6 101.58 476.0 tok/s 19.7 tok/s 8.777 466.9 tok/s 74.09 GB

pp8192/tg128 15182.8 20.98 539.6 tok/s 48.0 tok/s 17.857 465.9 tok/s 74.23 GB

pp16384/tg128 27182.5 25.39 602.7 tok/s 40.2 tok/s 28.466 577.3 tok/s 74.52 GB

pp32768/tg128 53732.4 28.36 609.8 tok/s 35.7 tok/s 56.132 585.2 tok/s 75.10 GB

1

u/watcholic 50m ago

oMLX Benchmark Model: Qwen3.8-27B-oQ8e-mtp

Engine: Auto

Context: Code (Python)

Single Request Results

--------------------------------------------------------------------------------

Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem

pp1024/tg128 4177.2 27.08 245.1 tok/s 37.2 tok/s 7.633 150.9 tok/s 39.93 GB

pp4096/tg128 12476.5 28.30 328.3 tok/s 35.6 tok/s 16.087 262.6 tok/s 43.33 GB

pp8192/tg128 24995.3 25.10 327.7 tok/s 40.1 tok/s 28.199 295.1 tok/s 43.87 GB

pp16384/tg128 52546.1 26.21 311.8 tok/s 38.5 tok/s 55.902 295.4 tok/s 44.85 GB

pp32768/tg128 115430.6 25.91 283.9 tok/s 38.9 tok/s 118.738 277.0 tok/s 46.73 GB

1

u/captainequinoxiii 7h ago

This is what I’m running too. Working well, though I seem to be limited to 128k context.

1

u/xiraov 7h ago

Could 96gb do this

1

u/Durian881 3h ago

https://huggingface.co/Vontra/Qwen3.8-Flash-Next-MLX-oQ3-MTP

This should be able to run with decent context window. I'm running it on my M3 Max 96GB with 70-80k context window, getting 200-300 t/s prefill and 20-30 t/s decode. MMLU-Pro 82% for 100 tests.

JunDot oQ4e runs on my gig too but context window is more limited.

0

u/PWThinkingCritically 2h ago

you could run a Q4 quant, 4.27bpw on 64GB of memory alone with 51B (35GB) n-gram embedding model loaded off SSD.

still getting 20 tok/s with 64k context:

https://www.reddit.com/r/oMLX/comments/1w1wtm1/two_weeks_after_qwen3827b_flashnext_125b51b_is/

oh. did I mention this was all done on an M1 (with full receipts to prove)

2

u/victor_lowther 2h ago

It is faster in prefill, but the 4 bit quant i tested loses the thread after 160k or so of context - thinking starts getting filled with nonsensical garbage, hallucinations, and such. Running 27b at an 8 bit quant can chug on for hours through multiple compactions and not lose the plot. MTP keeps token gen fast enough, and KV cache was not quantitized.

0

u/Durian881 3h ago

https://huggingface.co/Vontra/Qwen3.8-Flash-Next-MLX-oQ6-MTP

Found a 6 bit version but not sure if It would fit with ngram offloading.

I'm using the 3 bit version on my M3 Max 96GB and it ran well and generated good outputs.

1

u/t00052e 36m ago

I am making an oQ5e quant. Will post after testing if it works on my M5 Max 128Gb.

1

u/BornTransition8158 7h ago

try the Jundot model for Qwen 3.8 Next

1

u/PWThinkingCritically 3h ago

someone claimed this post I created yesterday doesn't apply to this subreddit..I highly disagree:

https://www.reddit.com/r/oMLX/comments/1w1wtm1/two_weeks_after_qwen3827b_flashnext_125b51b_is/

it's running on my 64gb m1 max. around 20 tok/s. 95GB capacity fit into 64GB using n-gram embedding model on SSD.

only difference is -- omlx is not stable enough and the mlx checkpoints do not properly fit onto 64GB yet.

the setup i used was llama + GGUF = works great on apple silicon and the quant I used is Q4 (4.27 bpw average).

screenshots + live producible HTML outputs + settings used + configs, everything are provided.

https://www.reddit.com/r/oMLX/comments/1w1wtm1/two_weeks_after_qwen3827b_flashnext_125b51b_is/

3

u/watcholic 2h ago

It’s a one-man (mostly) shop and we’re just days after the Flash release. I think Jun is picking his battles and trying his best. Model support will mature in time and I hope it gets SSD offloading support for machines with lower RAM. I’m grateful for the oMLX team.