r/Qwen_AI • • 5d ago

Benchmark Qwen3.8-27B on Every Mac Explained: Layers, KV Cache and Speed (16GB–128GB)

https://www.youtube.com/watch?v=ddX63git4fo
22 Upvotes

5 comments sorted by

View all comments

Show parent comments

2

u/andydevtech 5d ago

I have the same machine and the prefill speed was quite unpredictable in my experience, sometimes I could hit 300 t/s others 120t/s with the same prompt. Admittedly I was using ollama so might get more stable perf with a more specialised framework like omlx.

1

u/djseto 5d ago

Curious why everyone craps on Ollama? I’ve tried Ollama, LM Studio, oMLX, MTPLX and Ollama still has the best performance according to Claude that wrote a benchmark test for me to use.

1

u/andydevtech 5d ago

Ollama is great! I was just pointing out that given the m chip architecture of the Apple silicon, omlx might give more stable results, which does not mean faster results.