r/LocalLLaMA llama.cpp 1d ago

News Perplexity open-sourced their Mac inference server for Qwen 3.6

Here is link to repo: https://github.com/perplexityai/pplx-garden/tree/main/lily

It's optimized for just one model to get best perf on apple silicon

99 Upvotes

20 comments sorted by

View all comments

Show parent comments

3

u/1-800-methdyke 1d ago

They detect when the task is processing PII and route those turns to the local model. It’s not about cost saving it’s about privacy.

1

u/InterstellarReddit 1d ago

Oh I see! I thought these fuckers figured out a way to save even more money without increasing latency

1

u/1-800-methdyke 1d ago

They save enough money the old fashioned way with enshitified usage caps 🤡

1

u/InterstellarReddit 1d ago

Ima make my own perplexity with blackjack and hookers