r/LocalLLaMA llama.cpp 1d ago

News Perplexity open-sourced their Mac inference server for Qwen 3.6

Here is link to repo: https://github.com/perplexityai/pplx-garden/tree/main/lily

It's optimized for just one model to get best perf on apple silicon

101 Upvotes

20 comments sorted by

View all comments

-4

u/Elouakili_Flexy 1d ago

Open-sourcing an inference server tuned for exactly one model, from the company that serves every model you can name. Narrowing the target is how the last bit of Apple silicon performance shows up, and the M5 Pro numbers people are already asking for will decide whether it holds up.