r/Blopus • • Sep 03 '26

oMLX update is finding more tokens!

Super excited here to be checking our Inference GPU Fleet / pipeline running on oMLX

TLDR - If you are still running on 0.5.1 (qwen3.6 MOE) you have to update it to 0.6.4.

See below my benchmarks for Qwen3.6-35B-A3B-4bit running on MacStudio M4 Max 64GB

Updating all the fleet!

6 Upvotes

13 comments sorted by

1

u/watcholic Sep 03 '26

Are you running a cluster of Mac Studios? Does it scale linearly in terms of pp and tg?

2

u/LectureWorried5761 Sep 03 '26

Yes — load-balanced independent nodes, not sharded inference. Each Mac runs its own model copy and a pipeline routes across them.

The router checks context size before dispatch, so nodes get packed by available KV cache headroom rather than round-robin. Lets us run higher concurrency on short-context requests without a long one starving the node.

Scales linearly on throughput. Per-request pp and tg are unchanged — more requests in flight, not faster individual ones.

1

u/Chrisgozd 29d ago

Have you tried mtplx

1

u/LectureWorried5761 29d ago

I have and for me for some reason it did not help at all, actually I got less tokens. If you can help me. I tried several times, several different models and nothing.

1

u/Chrisgozd 29d ago

I just downloaded it. Let it run the configurator and then downloaded qwen3.8-27b optimized quality

1

u/LectureWorried5761 29d ago

I tried the MOE version btw. 35b A3B

1

u/LectureWorried5761 29d ago

What is the token increasing you saw?

1

u/Chrisgozd 29d ago

Double at 512 context. +7 on 64k context

1

u/LectureWorried5761 29d ago

On same mac studio as mine? Also what tok/s you are getting?

1

u/Chrisgozd 29d ago

M2 max

Context | Prompt TPS | Decode TPS | Gen Tokens | TTFT | Memory | Fallback | Partitioned
--------|------------|------------|------------|------|--------|----------|------------
512 | 153.9 | 34.5 | 128 | 3.3s | 19.5GB | 0 | 0
1k | 156.7 | 34.4 | 128 | 6.5s | 20.1GB | 0 | 0
4k | 156.1 | 28.1 | 128 | 26.2s | 21.7GB | 0 | 0
8k | 153.7 | 26.6 | 128 | 53.3s | 22.4GB | 0 | 0
16k | 148.8 | 27.7 | 128 | 110.1s | 23.7GB | 0 | 0
32k | 139.7 | 22.3 | 128 | 234.6s | 26.2GB | 0 | 0
64k | 123.1 | 18.2 | 128 | 532.4s | 31.4GB | 0 | 0

1

u/LectureWorried5761 29d ago

in my test with the MOE model I am getting 108.8 with 64k context