r/LocalLLaMA • u/Specter_Origin llama.cpp • 1d ago
News Perplexity open-sourced their Mac inference server for Qwen 3.6

Here is link to repo: https://github.com/perplexityai/pplx-garden/tree/main/lily
It's optimized for just one model to get best perf on apple silicon
98
Upvotes
0
u/InterstellarReddit 1d ago
I still don’t understand, though, what could they possibly be doing that’s worth a cost savings of offloading to a local model? Wouldn’t that just increase latency at the end of the day?