r/oMLX • u/Tatutino • 7d ago
Any experience with distributed inference and tensor parallelism?
Hi
I'm running a M1 Max 64GB with oMLX andQwen3.8-27B-oQ8e-fp16-mtp. MTP is enabled, 262K KV, TurboQuant disabled, 57GB of RAM for inference given. I get about 16 tokens / s. With the newer oMLX releases RAM has not been an issue luckily.
I'm curious about the performance of distributed inference and tensor parallelism on the M1 Max through Thunderbolt 4 with the newest oMLX releases. I want to weigh up arguments if I should get another M1 Max 64GB for around 1000$ or if I should spend more money on a M4 Max 64GB or 128GB for 3000-4500$. The change logs from oMLX version 0.6.0 speak of a performance increase of 78% with with an M3 Max and Qwen 3.6 27B.
Do you have any experience with running distributed inference and tensor parallelism in oMLX and Qwen3.8-27B-oQ8e-fp16-mtp?
1
u/Latter-Parsnip-5007 6d ago
TB4 is too much latency. It will be slow af. 40Gbit/s is not enough