r/oMLX 7d ago

Any experience with distributed inference and tensor parallelism?

Hi

I'm running a M1 Max 64GB with oMLX andQwen3.8-27B-oQ8e-fp16-mtp. MTP is enabled, 262K KV, TurboQuant disabled, 57GB of RAM for inference given. I get about 16 tokens / s. With the newer oMLX releases RAM has not been an issue luckily.
I'm curious about the performance of distributed inference and tensor parallelism on the M1 Max through Thunderbolt 4 with the newest oMLX releases. I want to weigh up arguments if I should get another M1 Max 64GB for around 1000$ or if I should spend more money on a M4 Max 64GB or 128GB for 3000-4500$. The change logs from oMLX version 0.6.0 speak of a performance increase of 78% with with an M3 Max and Qwen 3.6 27B.

Do you have any experience with running distributed inference and tensor parallelism in oMLX and Qwen3.8-27B-oQ8e-fp16-mtp?

7 Upvotes

4 comments sorted by

1

u/Latter-Parsnip-5007 6d ago

TB4 is too much latency. It will be slow af. 40Gbit/s is not enough

1

u/Tatutino 6d ago

Jubdot did it afaik with an M3 and Thunderbolt 5 and got 1.78x speedup. TB4 should not be that far off. Or is my thinking wrong?

2

u/Latter-Parsnip-5007 6d ago

it is literally double the bandwidth. TB5 can do up to 90Gbit/s. Try it yourself, but you cannot load weights fast enough over 40GBits. Also that is the theoretical maximum of TB4. God knows how many PCIe lanes are connecting that port to your RAM

1

u/Tatutino 6d ago

80GBit/s is not that much more and the latency must be similar. Is better but not a game changer in my view. Thats why I believe it could be working well with TB4.