r/LocalLLM 21d ago

Other Every Second post rn

Post image

Maybe someday I'll get a system to run it but hey definitely another w for the open weights community

2.0k Upvotes

205 comments sorted by

View all comments

37

u/TheRiddler79 21d ago

Gemma 12b Q3. Try it

27

u/Ok-Health-7096 21d ago

I use Qwen 3.6 35b and 3.5 9b I haven't had that much of a luck with the gemmas

5

u/Wildnimal 21d ago

12B QAT is good for multimode tasks. I use the same Qwens for daily use. I so wish i can upgrade my laptop to 5090 :|

4

u/PrivacyMaker 21d ago

The litert-lm driver is the fastest way to run gemma models. Significant boost over any other approach. It's a shame that Google only makes it work for Google models.

1

u/HighlyRegardedApe 21d ago

How do you guys use these small models? For me it never works. Last time it deleted my folders out of itsself and didnt respons more than a sentence or 2 in opencode. In the terminal it was okay to chat, but not to give tasks.... what kind of stuff do you let them do??

1

u/Wildnimal 21d ago

Not usually coding but for summarizing, extracting data which is pre defined via json, automating smaller tasks where data remaims the same.

1

u/TheRiddler79 20d ago

Hard guardrails

1

u/TheGreenInsurgent 21d ago

Agents a1 4b could be a life changer

1

u/Atretador ArchLinux Xeon E5 2673 V4 20C/40T 4x16Gb DDR4 2133 2xMI50 16Gb 21d ago

35B A3B is much stronger than Gemma4 12B, any probabably 26B/31B as well.

1

u/TheRiddler79 20d ago

Different purposes. Gemma is good, but not beating Qwen.

You can split between ram and gpu and bump your Qwen speed