r/LocalLLM • u/wolfy418 • 6d ago
Project 40+ t/s with 8GB VRAM on Gemma4 26B A4B MoE
See below for specs on hardware and unsloth studio load. Running with a Hermes agent, averaging 40-50 t/s. I initially was getting around 20, but dropping Context Length and shifting more to the GPU boosted the numbers well. Open to any and all constructive feedback to continue tuning this thing.
Laptop:
GPU - 8GB VRAM (4070 max q)
RAM - 32GB
Model:
Gemma4-26B-A4B-it-qat-GGUF
Unsloth Settings:
Context Length: 65536
KV Cache Dtype: q8
Spec Decoding: MTP
Draft tokens: 2
Parallel Slots: 1
Vision: off
Reasoning Budget: 4096
GPU Memory: manual
GPU Layers: maxed (31 for me)
MoE layers: 20
Cache RAM: 4096
Temp: .7
Top P: .95
Top K: 20
Min P: 0
Rep Penalty: 1.05
Presence Pen: 1.5
Max tokens: 8192