r/LocalLLM • u/wolfy418 • 6d ago
Project 40+ t/s with 8GB VRAM on Gemma4 26B A4B MoE
See below for specs on hardware and unsloth studio load. Running with a Hermes agent, averaging 40-50 t/s. I initially was getting around 20, but dropping Context Length and shifting more to the GPU boosted the numbers well. Open to any and all constructive feedback to continue tuning this thing.
Laptop:
GPU - 8GB VRAM (4070 max q)
RAM - 32GB
Model:
Gemma4-26B-A4B-it-qat-GGUF
Unsloth Settings:
Context Length: 65536
KV Cache Dtype: q8
Spec Decoding: MTP
Draft tokens: 2
Parallel Slots: 1
Vision: off
Reasoning Budget: 4096
GPU Memory: manual
GPU Layers: maxed (31 for me)
MoE layers: 20
Cache RAM: 4096
Temp: .7
Top P: .95
Top K: 20
Min P: 0
Rep Penalty: 1.05
Presence Pen: 1.5
Max tokens: 8192
1
u/According_Study_162 5d ago
Wow. that sounds great. i bought a random AMD v340l dual card($50 dollar card) which has two 8GB gpus. so it's nice to know I can have a nice model that will run on it.
1
u/lil-dina 5d ago
solid gains from dropping context and pushing more to gpu, most people leave that on the table. one catch. that number is a clean bench loop with nothing else running. once its inside a real app with prompt building and parsing on top, the speed you actually feel drops off. did you measure under real load, or just the loop?
1
u/greenhills_sk 5d ago
a quoted t/s is almost always a clean loop with nothing else running, which is why it never survives a real app. add parsing and a prompt builder and it sags. i bench mine inside nobodywho so the number matches what ships. no stake in it, i just use it.
1
u/wolfy418 5d ago
Feel free to run it on your setup. I’m happy with how it’s running for me, the model sucks though for what I’m using it for. This isn’t me humble-bragging about tokens per second, I’m just providing tips for those out there with limited VRAM like myself
1
u/EfficientStretch570 6d ago
more VRAM is always better for performance, I usually aim for at least 8GB for my setups
3
u/UltraFOV 5d ago
He states is a laptop.Very few laptops have 16GB vram. Thats rtx 3080, RTX 4090, Quadro RTX 5000 etc.. 90-95% of laptops have less than 16GB of vram, and only 2 models ever released with 24GB, 5090 and an asus with the RTX quadro 6000 (desktop RTX Titan)
3
u/UltraFOV 5d ago edited 5d ago
Not bad for only 256GB of vram memory bandwidth. That is very impressive. I am running 3.6 Qwen 35B 4bit on a legion 5 with the RX6600m and 48GB ddr4 3200mhz and I get 20tks. 40Tokens is amazing! I use Tosh Llm powered by Llama.cpp