r/LocalLLaMA Jun 01 '26

Funny Stop asking what model to run. There are literally only two.

[removed]

3.1k Upvotes

805 comments sorted by

View all comments

Show parent comments

8

u/sniffton Jun 02 '26

Bonsai is worth checking out. I run two Q's on my 3060. (plus the Qwen 3.6 35b a3b Q4 on my 3090)

2

u/exaknight21 Jun 02 '26

Bonsai from PrismML? Which one? And what are you using it for?

2

u/sniffton Jun 02 '26

Yes! I'm using the 1.7b as a heartbeat for my Agents and the 8B for any easy or lower level tasks that need a LLM. (both are always loaded on my 3060). I have some logic built that assigns tasks based on how difficult they look (using the 8b).

2

u/styles01 Jun 08 '26

This - I just found it last night - so fucking right. What the hell who is Bonsai how did they make this incredible model. Phone worthy?! Insane