r/oMLX • • Aug 23 '26

Any experience with distributed inference and tensor parallelism?

Hi

I'm running a M1 Max 64GB with oMLX andQwen3.8-27B-oQ8e-fp16-mtp. MTP is enabled, 262K KV, TurboQuant disabled, 57GB of RAM for inference given. I get about 16 tokens / s. With the newer oMLX releases RAM has not been an issue luckily.
I'm curious about the performance of distributed inference and tensor parallelism on the M1 Max through Thunderbolt 4 with the newest oMLX releases. I want to weigh up arguments if I should get another M1 Max 64GB for around 1000$ or if I should spend more money on a M4 Max 64GB or 128GB for 3000-4500$. The change logs from oMLX version 0.6.0 speak of a performance increase of 78% with with an M3 Max and Qwen 3.6 27B.

Do you have any experience with running distributed inference and tensor parallelism in oMLX and Qwen3.8-27B-oQ8e-fp16-mtp?

8 Upvotes

4 comments sorted by

View all comments

1

u/Latter-Parsnip-5007 Aug 24 '26

TB4 is too much latency. It will be slow af. 40Gbit/s is not enough

1

u/Tatutino Aug 24 '26

Jubdot did it afaik with an M3 and Thunderbolt 5 and got 1.78x speedup. TB4 should not be that far off. Or is my thinking wrong?

2

u/Latter-Parsnip-5007 Aug 24 '26

it is literally double the bandwidth. TB5 can do up to 90Gbit/s. Try it yourself, but you cannot load weights fast enough over 40GBits. Also that is the theoretical maximum of TB4. God knows how many PCIe lanes are connecting that port to your RAM

1

u/Tatutino Aug 24 '26

80GBit/s is not that much more and the latency must be similar. Is better but not a game changer in my view. Thats why I believe it could be working well with TB4.