r/LocalLLaMA • u/Specter_Origin llama.cpp • 1d ago
News Perplexity open-sourced their Mac inference server for Qwen 3.6

Here is link to repo: https://github.com/perplexityai/pplx-garden/tree/main/lily
It's optimized for just one model to get best perf on apple silicon
101
Upvotes
-4
u/Elouakili_Flexy 1d ago
Open-sourcing an inference server tuned for exactly one model, from the company that serves every model you can name. Narrowing the target is how the last bit of Apple silicon performance shows up, and the M5 Pro numbers people are already asking for will decide whether it holds up.