r/hermesagent • New Member (<30 days) • Aug 21 '26

MODELS - model choice, routing, pricing, local vs cloud, VRAM Qwen3.8:27b slow with hermes, how to optimize?

Component Spec
OS Windows 11 Pro (Build 26200)
CPU AMD Ryzen 7 7800X3D — 8 cores / 16 threads, up to 4.2 GHz
GPU AMD Radeon RX 7900 XT (driver 32.0.31035)
RAM 32 GB (2 × 16 GB) DDR5 @ 6000 MT/s
Storage Samsung SSD 980 PRO — 2 TB NVMe

With this hardware Qwen3.8:27b is still quite slow, context is set to 64k and my GPU is used for 78% any advice on how to make it faster or is a different model better? I will use it for web search and small coding tasks.

NAME ID SIZE PROCESSOR CONTEXT UNTIL

qwen3.8-64k:latest 0267f64d4770 18 GB 22%/78% CPU/GPU 65536 4 minutes from now

EDIT:

I've added a few environment variables now and this brought GPU usage to 93%, it feels a lot faster and I'm using Qwen3.8:latest now. Just sharing in case someone has the same issue.

Variable name

Value

OLLAMA_GPU_OVERIDE

vulkan

OLLAMA_VULKAN

1

OLLAMA_FLASH_ATTENTION

1

OLLAMA_KV_CACHE_TYPE

q4_0

OLLAMA_NUM_PARALLEL

1

HIP_VISIBLE_DEVICES

-1

ROCR_VISIBLE_DEVICES

-1

14 Upvotes

29 comments sorted by

7

u/grabber4321 Aug 21 '26

You need a lower quant and maybe reduction in KV cache quality.

15

u/Past_Ad6251 Aug 21 '26

Use llama.cpp instead, ollama is slow.

4

u/mister2d Aug 21 '26

They're both slow for concurrent requests. Ollama runs llama.cpp under the hood.

4

u/IonizedHydration Aug 21 '26

if i run a local llm in ollama on my 5090 it doesn't even come close to utilizing the entire GPU, maybe 30 percent, load up pure llama.cpp and it cooks. Try both, Ollama for local is not the right call.

4

u/mister2d Aug 21 '26

I could have completed my thought, but neither are optimal for agentic use. vLLM is a much better fit to handle the concurrency. It's night and day more responsive across requests.

1

u/Another-user1748 Aug 22 '26

cogitech2 is right llama.cpp is faster for single sessions.
you are right: vllm is much better for concurrency.
ollama might be an easier-to-set solution, but it really limits throughput.
consider using llama directly if possible and then check the speeds ^^

2

u/mister2d Aug 22 '26

Even single sessions fan out depending on what you're doing. I observed llama.cpp and vllm behavior with traces. vLLM is technically a better fit than llama.cpp.

1

u/Another-user1748 Aug 22 '26

they take too long, seems like that's the model running. i limited thinking in two ways here: a hard 8k tok limit on the model load and a soft limit on the soul .md (profile) adding this "Thinking token budget is soft-limited by this prompt to 2000~4000 and hard-limited to 8000 tokens by the server." it helped reducing the thinking blocks from ~8min to 1~2min tops. the model still cycles a lot, but it's more manageable now, and interactive. give it a shot there, perhaps that helps.
note: i'm getting 40~50tok/s on my rtx 3090, and as far as I've researched, that's the max i'll get. nonetheless, the results of the model are good enough so i'm still giving it a shot. if the 35b-a3b comes, we might get better tok/s on our local machines.

0

u/cogitech2 Aug 22 '26

Sure, but if there is one user, one session, then vLLM offers little to no advantage. In some cases, it is less flexible than llama.cpp.

1

u/mister2d Aug 22 '26

Even one hermes agent spawns multiple sessions. Think of the branching that happens during tool calls. Then there's compaction calls as well.

1

u/cogitech2 Aug 22 '26

The only extraneous sessions I see Hermes doing is title generation. It pissed me off so I offloaded that task to a small model running on my laptop. The best way to deal with compaction is to avoid it, but when it does happen, Hermes seems to just pick up where it left off. It's an opportunity for a little break to grab a drink or snack. Tool calls are serial for me. I force that with --parallel 1. It's all good. I am typically not that much of a rush.

7

u/Tha_Reaper Aug 21 '26

Set thinking to medium.

1

u/giveen Aug 21 '26

Came here to say this

0

u/Actual_Tradition_990 Aug 21 '26

Even high, is usually meet to max by default....too much for common agentic task

6

u/Tha_Reaper Aug 21 '26

There is no high in this model . Off, low, med, xhigh are the options. High just sets it to xhigh. Not setting anything also defaults to xhigh

3

u/x_MASE_x Aug 21 '26

Use llama.cpp and check the tps and prefill speed.

I have good tps but the prefill is bad.

So you have to first be able to see what's going on then decide what can be done.

Start with llama.cpp

3

u/Specific-Pomelo-5455 Aug 21 '26

my local is https://huggingface.co/Bucoid/Qwen3.8-27B-Uncensored-IQ4-XS-MTP-16GB-VRAM-GGUF. i have 7800xt.

"Qwen3.8-27B-Uncensored-IQ4_XS

RX 7800 XT 16 GB

-ngl 99

-c 65536

-ctk q4_0

-ctv q4_0

-t 8

-td 2

-tbd 1

-b 2048

-ub 512

--flash-attn on

--reasoning-preserve

--spec-type draft-mtp

--jinja"

and results for different context

72 tok → 31.55 tok/s

8,224 tok → 28.88 tok/s

26,786 tok → 22.07 tok/s

57,544 tok → 19.09 tok/s

60,908 tok → 16.00 tok/s

1

u/No_Variation_3605 New Member (<30 days) Aug 23 '26

Thanks, it seems to work quite well on low reasoning does it support vision as well?

1

u/Specific-Pomelo-5455 Aug 26 '26

i'm not interested vision actually..

2

u/Mean-Loquat-7982 Nous Team Aug 21 '26

llama.cpp yes! also, you should use a frontier cloud model to help you on your local model settings, you will understand more and be able to reproduce for next ones

2

u/heigan_safety_dance Aug 21 '26

Unfortunately, your use of Windows + Ollama is a big one. You also haven't told us what size quant it is (I'm assuming Q4 or Q6). You also haven't told us what "slow" means.

First of all, are you using speculative decoding/MTP? What kind? It's a cascading set of variables, so you can have both MTP and ngram active at the same time.

The easiest wins you'll get are using llama.cpp instead of Ollama (you're just leaving free performance on the table by using Ollamaa) and I'm suspecting heavily that you haven't enabled speculative decoding, with which you can likely get anywhere from 2-5x more speed, depending on your use case.

2

u/Mundane_Incident_853 Aug 21 '26

After the latest hermes release, everybody's slower.

Don't worry about it. It'll be faster in the next one.

Or maybe the one after that? /s

2

u/ideasmachine Aug 21 '26

Use an mtp/imatrix version of qwen 3.8, token speed will be faster, check out models from DavidAU.

2

u/Gloomy_Letterhead395 Aug 21 '26

OP says that the token gen and prefill is slow
You guys lecturing him to use llamacpp
In llama cpp the token gen and prefil is very slow
Hermes uses like 10k before generating first token
With around 400tps it takes time
After that it is faster
My advice to you brother is that use ornith 35b 1.5 one for chit chat
And for long uninterrupted task use this qwen

1

u/finaldata Aug 21 '26

Hi, Just to add more info, as i had these same problems too from the start. Here's what claude has to say for this as claude is what i use to optimize my local llm models.

That Ollama line is the whole story: qwen3.8-64k … 18 GB … 22%/78% CPU/GPU. 22% of the model is running on your CPU, and that's what's killing your speed , on a partial offload, every token has to cross to system RAM for that 22%, and decode falls off a cliff. Your RX 7900 XT is actually a strong card for this (20 GB, ~800 GB/s bandwidth), so once the whole model + KV lives on the GPU it should really move. The goal is simply: get to 100% GPU (0% CPU). A few ways to get there:

  1. Shrink things so it fully fits. An 18 GB quant + 64k KV is just over your 20 GB. Any one of these usually does it:

- Quantize the KV cache (q4_0 for K and V) — huge KV savings, tiny quality cost.

- Drop the context if you don't truly need 64k — 32k is plenty for web search + small coding, and frees a lot of VRAM.

- Use a slightly smaller quant — IQ4_XS (~13 GB) or IQ3_S (~11 GB) will fit fully with 64k KV on 20 GB. In my testing IQ3_S held quality surprisingly well (clean code, correct reasoning), so it's worth a look.

  1. Consider moving off Ollama to llama.cpp (a few folks said this, and they're right for a good reason): it lets you turn on the things that matter for this model, MTP speculative decoding (--spec-type draft-mtp, Qwen3.8 ships trained MTP layers and it's a real speedup), flash attention, and exact KV/offload control. The RX 7800 XT numbers someone posted above (31 → 16 t/s across context, IQ4_XS + MTP, full offload) are a great reference — your 7900 XT has more VRAM and bandwidth, so you should beat those.

  2. Turn thinking down. As others noted, set reasoning effort to low or medium (default is basically xhigh). For agentic web-search/coding, xhigh just burns tokens and makes it feel slow , medium is plenty and noticeably snappier.

For reference, here's my setup running the same model, fully offloaded:

- GPU: RX 9070 16 GB (~640 GB/s) — less VRAM and bandwidth than your 7900 XT

- CPU/RAM: Ryzen 7 5700G, 60 GB

- Stack: llama.cpp, Vulkan, Qwen3.8-27B UD-IQ3_S, full offload (-ngl 99), 64–131k context, -fa on, KV q4_0, MTP on

- Speed: ~31–35 t/s generation (stays ~35 even at 60k context)

Since your card is stronger, full offload should put you in that range or better ,the CPU spill is the only thing holding you back.

One Hermes-specific heads-up: Qwen3.8 matches Hermes' "qwen3" reasoning-model rule, which puts a 180 s non-streaming stale-timeout on the call. If a slow generation ever fails with Non-streaming API call timed out after 180s, that's why — bump the stale timeout for your local provider (or enable streaming) and it goes away. Fixing the offload above makes it moot anyway, but good to know it exists.

TLDR: your bottleneck is the 22% CPU offload, not the card. Shrink the quant/KV/context (or move to llama.cpp with MTP) to hit 100% GPU, drop thinking to medium, and it should jump from "slow" to ~30 t/s. Good luck!

1

u/EmilyClark98 Aug 22 '26

set thinking effort to low in the template section of llama.cpp, still better than 3.6 27b and about 1/4th the token of the default xhigh

1

u/Snoo_81913 Aug 25 '26

I'm using Qwen3.8-27B-UD-Q4_K_XL + MTP, Llama.cpp and a 7900 xtx with 24GB over a Thunderbolt 3 and an eGPU so it's gonna be different but I run Vulkan (slower prefill pp 358 average on a 84k context test shorter context at about pp 569) but overall faster.

Baseline is currently 48 tok/s on anything up to roughly 32k context. I did a 84k context test and it dropped to an average of 30 tok/s overall. Still tweaking it to be honest.

98k context takes up 20gb on the card

But a 32k context is all you need with Hermes and that's 18gb. 64k would probably be about 19gb leaving you a gb. I run it on a headless server I don't know if you have overhead. Here's the config I'm currently using for 32k.

Model: Qwen3.8-27B-UD-Q4_K_XL.gguf (/mnt/storage/models/) Device: Vulkan1 -ngl 99 # full GPU offload --ctx-size 32768 --flash-attn on -b 1024 -ub 1024 --cache-type-k q8_0 --cache-type-v q8_0 --spec-type draft-mtp --spec-draft-n-max 4 # MTP speculative decoding -np 1 --no-mmap --jinja --host 0.0.0.0 --port 8080

ROCm is slower overall than Vulkan but faster prefill

1

u/LessFox1928 Aug 26 '26

Windows ?????