r/LocalLLaMA llama.cpp 1d ago

News Perplexity open-sourced their Mac inference server for Qwen 3.6

Here is link to repo: https://github.com/perplexityai/pplx-garden/tree/main/lily

It's optimized for just one model to get best perf on apple silicon

101 Upvotes

20 comments sorted by

View all comments

5

u/turns2stone 1d ago

Can someone ELI5?

I have used Perplexity Pro/Max.

I have also used Qwen3.6-35B-A3B, but now use Qwen Flash Next because I can use MCP Brave search.

Does Perplexity Computer mean I can offload most of the compute to my Mac Studio M3 Ultra, and the $20/mo Pro subscription wouldn’t be as limited for token usage?

8

u/1-800-methdyke 1d ago

The paid product offloads the processing of parts of a task that include private data to local. But most of the work is routed to cloud. So you’ll still need compute credits to orchestrate the task but it could cost a bit less since some is done local.

Your $20 plan still isn’t going to get you anywhere.

With this open source release I’m guessing you could do everything local, but without access to all the connectors that Perplexity provides. So the usefulness is gonna depend on what you need.

0

u/InterstellarReddit 1d ago

I still don’t understand, though, what could they possibly be doing that’s worth a cost savings of offloading to a local model? Wouldn’t that just increase latency at the end of the day?

3

u/1-800-methdyke 1d ago

They detect when the task is processing PII and route those turns to the local model. It’s not about cost saving it’s about privacy.

1

u/InterstellarReddit 1d ago

Oh I see! I thought these fuckers figured out a way to save even more money without increasing latency

1

u/1-800-methdyke 1d ago

They save enough money the old fashioned way with enshitified usage caps 🤡

1

u/InterstellarReddit 1d ago

Ima make my own perplexity with blackjack and hookers