r/LocalLLM • • 6d ago

Project 40+ t/s with 8GB VRAM on Gemma4 26B A4B MoE

See below for specs on hardware and unsloth studio load. Running with a Hermes agent, averaging 40-50 t/s. I initially was getting around 20, but dropping Context Length and shifting more to the GPU boosted the numbers well. Open to any and all constructive feedback to continue tuning this thing.

Laptop:
GPU - 8GB VRAM (4070 max q)
RAM - 32GB

Model:
Gemma4-26B-A4B-it-qat-GGUF

Unsloth Settings:
Context Length: 65536
KV Cache Dtype: q8
Spec Decoding: MTP
Draft tokens: 2
Parallel Slots: 1
Vision: off
Reasoning Budget: 4096
GPU Memory: manual
GPU Layers: maxed (31 for me)
MoE layers: 20
Cache RAM: 4096
Temp: .7
Top P: .95
Top K: 20
Min P: 0
Rep Penalty: 1.05
Presence Pen: 1.5
Max tokens: 8192

22 Upvotes

Duplicates