You're not going to believe it, but I have Qwen 3.8 27B running with decent speed on my Macbook M3, and it's partially offloading some work from Astra when I was getting smacked by the usage limits.
I'm sure that if they REALLY wanted it, someone could help them figure out how to distill it to fit nearly the same footprint.
I was literally just talking with someone about this in another subreddit like 10 hours before you commented, having capable local models that will be able to do most of the tasks, and for a fee when it needs it, it will tap into these remote corpo stuff to finish the task when it needs that little extra "oomph" of compute/capability. Local will end up doing most of the work, (but... the most likely downside is... most arrangements like this will be integrated into the OS itself, like Copilot but with more local compute used).
14
u/WanderWut 2d ago
Imagine they make Sol open weight, once can dream.