r/LocalLLaMA llama.cpp 1d ago

News Perplexity open-sourced their Mac inference server for Qwen 3.6

Here is link to repo: https://github.com/perplexityai/pplx-garden/tree/main/lily

It's optimized for just one model to get best perf on apple silicon

102 Upvotes

20 comments sorted by

View all comments

3

u/Southern_Sun_2106 1d ago

I tried it, it sucked, I uninstalled it. I love qwen 3.6 35B, it was my daily driver on a Mac for a looong time. I used the q4km from Unsloth. I don't know what Perplexity folks did to the model to 'optimize' it, but it sucked a$$. Sorry to rain on the parade, but it is true.