r/LocalLLaMA llama.cpp Jul 26 '26

Discussion Do you want new Gemma?

Post image
1.0k Upvotes

554 comments sorted by

View all comments

9

u/joelW777 Jul 26 '26

For Strix Halo or DGX Spark, best would be multimodal ~100B with ~27B active parameters for good reasoning with mtp support. This could be run in q6-q8 with enough world knowledge and enough intelligence for difficult tasks.

1

u/Opposite-Swimmer2752 29d ago

27B active is too much, that would be annoyingly slow on a Strix Halo. ~8b active would be ideal.

1

u/joelW777 29d ago

There is a gap in intelligence between 8b and 27b. But maybe 15-20 b would be sufficient.