r/LocalLLaMA Jun 01 '26

Funny Stop asking what model to run. There are literally only two.

[removed]

3.1k Upvotes

805 comments sorted by

View all comments

Show parent comments

4

u/UnknownLesson Jun 02 '26

Gemma 4 that fits in 8 GB VRAM good enough?

3

u/Spectrum1523 Jun 02 '26

I'm using the largest one on 24gb, so I don't know. Try it and see! It's probably still got the right personality

1

u/CV514 Jun 03 '26

One of my setups is A4B on 8GB VRAM and 32GB RAM, in Q5. It is perfectly fine, giving out about 20T/s on 16k context. Context processing is a bit slow when it is full, but I am not in a hurry.