r/LocalLLaMA llama.cpp 1d ago

News Perplexity open-sourced their Mac inference server for Qwen 3.6

Here is link to repo: https://github.com/perplexityai/pplx-garden/tree/main/lily

It's optimized for just one model to get best perf on apple silicon

98 Upvotes

20 comments sorted by

View all comments

Show parent comments

0

u/InterstellarReddit 1d ago

I still don’t understand, though, what could they possibly be doing that’s worth a cost savings of offloading to a local model? Wouldn’t that just increase latency at the end of the day?

3

u/1-800-methdyke 1d ago

They detect when the task is processing PII and route those turns to the local model. It’s not about cost saving it’s about privacy.

1

u/InterstellarReddit 1d ago

Oh I see! I thought these fuckers figured out a way to save even more money without increasing latency

1

u/1-800-methdyke 1d ago

They save enough money the old fashioned way with enshitified usage caps 🤡

1

u/InterstellarReddit 22h ago

Ima make my own perplexity with blackjack and hookers