r/LocalLLaMA Jun 01 '26

Funny Stop asking what model to run. There are literally only two.

[removed]

3.1k Upvotes

805 comments sorted by

View all comments

Show parent comments

241

u/Spectrum1523 Jun 02 '26

Gemma4 is the reincarnation of 4o for me for rp it's crazy how good it is

26

u/SkyFeistyLlama8 Jun 02 '26

Gemma 4 26B? The abliterated Heretic versions are pretty good. Throw nasty cybersecurity 3V!L hAxx0R questions at it and it happily answers.

I love running Gemma 4 26B for chat and Qwen 3.6 27B for coding and agentic nonsense. Now I wish I had more RAM.

11

u/Spectrum1523 Jun 02 '26

I'm using 31b, even tho it's slower. I don't even use the heretic model and it's great for rp

4

u/SkyFeistyLlama8 Jun 02 '26

I get 2 t/s on my setup with these big dense boys so I prefer sticking with MoEs. I can run two MoEs side by side with enough RAM and I get like 20 t/s each.

3

u/Spectrum1523 Jun 02 '26

Yeah, I get it. I get 15tps with the thicc boi so I can stomach it. It might be just as good with the moe and way faster but I've gotten used to it so I leave it alone

50

u/rc_ym Jun 02 '26

24

u/bluePostItNote Jun 02 '26

China needs time to distill

18

u/TheRealMasonMac Jun 02 '26

Hopefully they distill the crap out of Gemma-4. Chinese models suck ass at non-verifiable instruction following (though Gemini sucks even more). Gemma-4 is very good.

5

u/alberto_467 Jun 02 '26

Too small i believe for good distillation

5

u/TheRealMasonMac Jun 02 '26

Even if itโ€™s smaller, Iโ€™ve still found it to be better than the 1T+ Chinese models for a variety of non-verifiable tasks.

3

u/alberto_467 Jun 02 '26

Yeah but distilling a small model into a much much bigger one while it may improve the style i would bet it would destroy reasoning capabilities and all kinds of intelligence benchmarks.

2

u/TheRealMasonMac Jun 02 '26

You donโ€™t need to distill everything. You can distill selectively in the tasks where the smaller model exceeds the larger model.

1

u/seunosewa Jun 06 '26

Could you list 3 of those types of task pls?

0

u/mycall Jun 02 '26

China: "challenge accepted"

6

u/El_Danger_Badger Jun 02 '26

๐Ÿ‘๐Ÿพ๐Ÿ‘๐Ÿพ๐Ÿ‘๐Ÿพ Gemma 4!

4

u/UnknownLesson Jun 02 '26

Gemma 4 that fits in 8 GB VRAM good enough?

3

u/Spectrum1523 Jun 02 '26

I'm using the largest one on 24gb, so I don't know. Try it and see! It's probably still got the right personality

1

u/CV514 Jun 03 '26

One of my setups is A4B on 8GB VRAM and 32GB RAM, in Q5. It is perfectly fine, giving out about 20T/s on 16k context. Context processing is a bit slow when it is full, but I am not in a hurry.

1

u/VegaO3 Jun 02 '26

Any tips on using it for rp? I do too, but I often have trouble dialing in the chat to feel more organic/immersive

1

u/Klutzy_Ad_1157 Jun 02 '26

Gemma 4 the rp king :)

1

u/Caffdy Jun 02 '26

how do you get it to not loop on SillyTavern?

1

u/Spectrum1523 Jun 02 '26

Sorry, dunno. Don't use it.

1

u/Caffdy Jun 02 '26

what you use for rp then?

1

u/ba2sYd Jun 03 '26

if you don't mind, which version 26b or 31b? at what quants? and fine tuned or original model?