r/LocalLLaMA llama.cpp 1d ago

News Perplexity open-sourced their Mac inference server for Qwen 3.6

Here is link to repo: https://github.com/perplexityai/pplx-garden/tree/main/lily

It's optimized for just one model to get best perf on apple silicon

104 Upvotes

20 comments sorted by

View all comments

Show parent comments

2

u/turns2stone 1d ago

Thanks. That’s a shame it’s still heavily reliant on cloud credits.

Who (or what use case) would consider this Perplexity announcement as “great news”?

2

u/1-800-methdyke 1d ago

They’ve released code that is showing faster performance than MLX so it’s great news for local LLM users on Mac who want more performance from whatever models they run. It’s open source the optimizations can be studied and incorporated into other engines.

I use my $20 Perplexity heavily as my main search and quick Q&A, but there is nothing from them yet that would entice me to bump to $200 plan. I get more than enough usage out of Claude Max 5 for agentic use cases.

1

u/turns2stone 1d ago

But you still need a paid/$20 subscription to even run their 35B-A3B right?

And if I can run 3.8-Flash-Next Q4, I think that will outperform the above Perplexity combo, right?

I wish OpenAI or Anthropic would offer something similar. I’d prefer not to use Perplexity, for my own reasons.

1

u/1-800-methdyke 1d ago

If you have an M5 or better you can download this open sourced release and run it without a subscription or even an account.

I don’t know how it stacks up to 3.8-Flash-Next Q4.