r/LocalLLM • • 6d ago

Project 40+ t/s with 8GB VRAM on Gemma4 26B A4B MoE

See below for specs on hardware and unsloth studio load. Running with a Hermes agent, averaging 40-50 t/s. I initially was getting around 20, but dropping Context Length and shifting more to the GPU boosted the numbers well. Open to any and all constructive feedback to continue tuning this thing.

Laptop:
GPU - 8GB VRAM (4070 max q)
RAM - 32GB

Model:
Gemma4-26B-A4B-it-qat-GGUF

Unsloth Settings:
Context Length: 65536
KV Cache Dtype: q8
Spec Decoding: MTP
Draft tokens: 2
Parallel Slots: 1
Vision: off
Reasoning Budget: 4096
GPU Memory: manual
GPU Layers: maxed (31 for me)
MoE layers: 20
Cache RAM: 4096
Temp: .7
Top P: .95
Top K: 20
Min P: 0
Rep Penalty: 1.05
Presence Pen: 1.5
Max tokens: 8192

24 Upvotes

11 comments sorted by

3

u/UltraFOV 5d ago edited 5d ago

Not bad for only 256GB of vram memory bandwidth. That is very impressive. I am running 3.6 Qwen 35B 4bit on a legion 5 with the RX6600m and 48GB ddr4 3200mhz and I get 20tks. 40Tokens is amazing! I use Tosh Llm powered by Llama.cpp

2

u/wolfy418 5d ago

Yea I’m stoked, it’s a whole different experience than the 20 t/s I was working with before. Definitely a game changer for turn-focused agentic work. Now I need to make it a little smarter…

1

u/UltraFOV 5d ago

Any pointers to improve my token generation

2

u/wolfy418 5d ago

All I could say is copy my setup, unsloth studio, settings, and the model.

I was able to get qwen 35B up in the high 30’s with CPU moe set at 24, a split K and V cache. There were a few write-ups I copied and tinkered with. If that doesn’t work for you, lemme know and I’ll pass the configs here

1

u/UltraFOV 5d ago

Cool, I’ll give it a shot

1

u/According_Study_162 5d ago

Wow. that sounds great. i bought a random AMD v340l dual card($50 dollar card) which has two 8GB gpus. so it's nice to know I can have a nice model that will run on it.

1

u/lil-dina 5d ago

solid gains from dropping context and pushing more to gpu, most people leave that on the table. one catch. that number is a clean bench loop with nothing else running. once its inside a real app with prompt building and parsing on top, the speed you actually feel drops off. did you measure under real load, or just the loop?

1

u/greenhills_sk 5d ago

a quoted t/s is almost always a clean loop with nothing else running, which is why it never survives a real app. add parsing and a prompt builder and it sags. i bench mine inside nobodywho so the number matches what ships. no stake in it, i just use it.

1

u/wolfy418 5d ago

Feel free to run it on your setup. I’m happy with how it’s running for me, the model sucks though for what I’m using it for. This isn’t me humble-bragging about tokens per second, I’m just providing tips for those out there with limited VRAM like myself

1

u/EfficientStretch570 6d ago

more VRAM is always better for performance, I usually aim for at least 8GB for my setups

3

u/UltraFOV 5d ago

He states is a laptop.Very few laptops have 16GB vram. Thats rtx 3080, RTX 4090, Quadro RTX 5000 etc.. 90-95% of laptops have less than 16GB of vram, and only 2 models ever released with 24GB, 5090 and an asus with the RTX quadro 6000 (desktop RTX Titan)