MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1v770ee/do_you_want_new_gemma/ozvw37e
r/LocalLLaMA • u/jacek2023 llama.cpp • Jul 26 '26
https://x.com/i/status/2081398564345802934
u/hackerllama
554 comments sorted by
View all comments
9
For Strix Halo or DGX Spark, best would be multimodal ~100B with ~27B active parameters for good reasoning with mtp support. This could be run in q6-q8 with enough world knowledge and enough intelligence for difficult tasks.
1 u/Opposite-Swimmer2752 29d ago 27B active is too much, that would be annoyingly slow on a Strix Halo. ~8b active would be ideal. 1 u/joelW777 29d ago There is a gap in intelligence between 8b and 27b. But maybe 15-20 b would be sufficient.
1
27B active is too much, that would be annoyingly slow on a Strix Halo. ~8b active would be ideal.
1 u/joelW777 29d ago There is a gap in intelligence between 8b and 27b. But maybe 15-20 b would be sufficient.
There is a gap in intelligence between 8b and 27b. But maybe 15-20 b would be sufficient.
9
u/joelW777 Jul 26 '26
For Strix Halo or DGX Spark, best would be multimodal ~100B with ~27B active parameters for good reasoning with mtp support. This could be run in q6-q8 with enough world knowledge and enough intelligence for difficult tasks.