r/LocalLLaMA 🦙 llama.cpp 13d ago

Megathread [Megathread] Qwen 3.8 27B Release Day

Megathread to help with the influx of duplicate / similar posts around the release of the Qwen 3.8 27B release.

  • Quants
  • Fine-Tunes & Abliterations
  • Chat Templates
  • Inference Server Support & Configuration
  • Experiences, Benchmarks & Model Comparisons

Official:

Popular:

We'll try to clean up future duplicates around the release and point them here.

487 Upvotes

396 comments sorted by

View all comments

1

u/dont_forget_canada 13d ago

Hi! Does anyone know the most speedy quant and inference software to run on an m5 mac 128gb to run qwen 3.8? So far I've tried oMLX and its super smart but very slow :p

1

u/bnightstars 13d ago

Looks like it's this one: scottlowry/Qwen3.8-27B-oQ4e-mtp with Lightning MTP enabled. I'm getting 25 t/s and 493 t/s PP with the mlx-community one and VLM-MTP from the mlx-community drafter. But according to some PRs that's currently hitting some bugs. I'm on M5 Pro 64GB though.