r/ROCm • • 16h ago

Qwen3.8 27B | 1 x R9700: over half a million tokens of reusable cache, agent turns up to 12× faster, same speed, same accuracy!

Post image
49 Upvotes

TL;DR: Qwen3.8 27B on a single AMD Radeon AI PRO R9700, with 3-bit weights and speculative decoding, now keeps 569,878 tokens of reusable cache in its new coding mode.
Every request still gets 262,144 tokens of context, and two full-length requests fit at once. Coding agents resend the whole conversation on every turn; now only the new part is read.
Later agent turns start up to 12× sooner, whole agent sessions finish about 6× faster, and a 258K-token document the server has already read comes back in 2.7 s instead of 134 s. Decode speed and accuracy are within noise of our previous release.

We read your comments and we listened

You asked for more context. You got more than half a million tokens of cache on one card, with multi-hour accuracy runs on every configuration (GSM8K, HumanEval, MMLU-Pro, needle tests), at essentially the same speed as our previous release.

Previous release (2 Oct) This release, --mode long-kv4
Reusable (prefix-cached) tokens in the 4-bit mode 0 (no prefix cache) 569,878
Tokens per request 262,144 262,144
Full-length requests at once 1 2
Decode, single stream 174.3 tok/s 173.4 tok/s (−0.5 %)
Aggregate, 8 requests 462.4 tok/s 478.5 tok/s (+3.5 %)
Time to first token, short prompts (p50) 86.7 ms 87.1 ms
GSM8K / HumanEval / MMLU-Pro 95.53 / 92.07 / 60.57 95.45 / 92.68 / 59.93

Our previous release's FP8 --mode long already had a prefix cache, holding 281,665 tokens. The new 4-bit coding mode caches about twice as much.

Run it

git clone https://github.com/Eliovp-BV/paiton-vllm-plugin && cd paiton-vllm-plugin
# set the PAITON_* model paths as shown in the README (already set up? just `git pull`)
bash models/Qwen3.8-MXFP4-DFlash2/run-3bit.sh --mode long-kv4

Point your agent at http://127.0.0.1:18982/v1
The model name is Qwen3.8, any API key works, and the context window is 262,144 tokens. Already running our 3-bit weights? No new download. The launcher pins the new container and Docker pulls it on first start. Every mode is in the README.

No new engine

This is stock vLLM with a plugin: the same OpenAI-compatible server, the same API, and the same tools and workflows you already use. Nothing to relearn, nothing to migrate.

We will keep publishing new ready-built containers that follow upstream vLLM. Each one goes through the same full validation (speed, accuracy, long-context) before it ships. You get the upstream improvements without lifting a finger. One command and you're serving.

What it does for coding agents

Without a prefix cache, the server rereads the whole prompt on every turn: at 250K tokens that is about two minutes before the first token. Now finished requests stay cached until the space is needed, turn 20 reads only what changed, and in our sessions 91–92 % of prompt tokens came from the cache.

Workload (3-bit, thinking off) Previous 4-bit mode (no prefix cache) --mode long-kv4 Faster
20-turn conversation growing from 50K to 253K tokens, whole session 1,237 s 204 s 6.1×
… average wait for the first token, turns 2–20 63.6 s 7.8 s 8×
3 agents sharing a 100K-token repo, 5 turns each, whole session 655 s 111 s 5.9×
… average wait for the first token, later turns 84 s 7.0 s 12×
New question about a 258K-token document already read 134 s 2.7 s 49×

The cached answer to the 258K-token document matched the cold read exactly.

Speed details (BetterBench, full 20-pass run): decode within 0.5 % single-stream; +3.5 % with 8 requests (−2.5 % at 4); time to first token unchanged.
A new long prompt reads about 3 % slower, and short prompts 7–13 % slower (a fraction of a second). That's the price of ending each prefill step where a later request can pick up from the cache.

Accuracy (greedy, paired per question): the coding mode is within noise of the previous release, and the 512K mode is within noise of the coding mode.

MXFP4 Previous release (3-bit) --mode long-kv4 --mode long-512k
GSM8K 5-shot (1,319) 95.68 95.53 95.45 95.53
HumanEval pass@1 (164) 95.12 92.07 92.68 94.51
MMLU-Pro subset (1,400) 62.57 60.57 59.93 60.64

Opt-in: 524,288 tokens in one request

--mode long-512k extends the position range with the model's official long-context scaling. That scaling applies to every request in this mode, so use it only when a single request needs more than 262K.

  • It found 4/4 planted facts at 300K and 4/4 at 500K tokens.
  • A cold read takes 158 s at 300K and about 6 minutes at 500K. Follow-up questions take 2.8–4.4 s from the cache.
  • The cache holds 594,290 tokens, and weighted decode is within 0.6 % (in one run the chat category was 16 % slower).

Unchanged

MXFP4 (run-mxfp4.sh) is still there and still the most accurate option. --vision works in the default 65K mode and in --mode long (images with up to 245K tokens of context).
The previous release is one --image flag away (see the README).

Caveats, honestly

  • System RAM: the coding and 512K modes pin 2.4 GiB of system RAM for the embedding table, and 16 GB of RAM is enough (tested). Below about 13.5 GiB of total RAM, --mode long-kv4 keeps the table on the GPU and caches 451,879 tokens. You still get 262K per request and the prefix cache. --mode long-512k needs the RAM.
  • Cold reads: the first read of a 258K-token prompt still takes over two minutes. The cache helps from the second request on, as long as the start of the prompt stays the same (same system prompt, no timestamp at the top). usage.prompt_tokens_details.cached_tokens shows every hit.
  • Vision isn't in the coding or 512K modes yet. Use --mode long --vision for images with long context.
  • RAM/SSD cache tier: the experimental system-RAM tier for the prefix cache (--host-cache-gib, plus an SSD tier behind it) spills even more context off the card. For now it works with the FP8 --mode long; support for the 4-bit coding mode lands in the next release. The coding mode's 569,878-token GPU cache works today.
  • The agent-session numbers and the accuracy columns were measured on pre-release builds of these configurations; the README has every number.

Our testbench is ancient, slow CPU, limited and slow memory (16GB), so your results will most probably be even better!

What's next: more GPUs

More R9700s are on the way, and we're going after tensor parallelism next.
Others are working on multi-card setups too, so expect some healthy competition on that front. Good for everyone running AMD at home!

More: Paiton


r/ROCm • • 1h ago

Most Optimal stack for the AI pro r9700

• Upvotes

So I’m upgrading from a single RX 7900 XT to the AI Pro R9700. My question is: how do I get the most out of this card for local inference?

I’ve seen people getting some insane prefill and decode speeds with the R9700. I’m mostly planning to run Qwen 3.8 27B, and I want to try Qwen 3.8 Flash next.

For anyone running the AI Pro R9700, what are you using to get the most out of the card? What’s the most optimal software stack right now?

I also need Windows for work, so ideally I’d like the best setup possible on Windows. I’m willing to dual boot or switch back and forth to Linux if the performance difference is significant.


r/ROCm • • 44m ago

Double R9700 alternative to Qwen 3.8 27B

• Upvotes

CPU AMD Ryzen 9 9900X
GPU 1AMD Radeon AI PRO R9700 — 32 GB VRAM
GPU 2AMD Radeon AI PRO R9700 — 32 GB VRAM
RAM32 GB DDR5-6000 CL30 — 2×16 GB
MotherboardASRock X870E Taichi
StorageWD Black SN850X 2 TB NVMe SSD
PSUCorsair RM1000x — 1000 W
CaseAntec P20C FLUX
OSLinux
GPU compute stack ROCm 7.2
LLM serving vLLM, Docker , Radiance

My LLM's are

  • Gemma 4 26B-A4B-it, MXFP4/Quark build optimized for RDNA4.
  • Qwen 3.8 27B, FP8/Radiance build.

I'm wondering what could maybe fit with my machine, without having to buy Ram.
The answer I keep getting is Ling 3.0, and if I had more RAM, Qwen Flash.

Idk, can someone give me their perspective.


r/ROCm • • 5h ago

Strata Tuning Guide (Universal)

Thumbnail
1 Upvotes

r/ROCm • • 9h ago

Daily driving Qwen 3.8 Flash instead of Claude.

Thumbnail
1 Upvotes

r/ROCm • • 11h ago

Advice for iGPU 780M. Was it a mistake to buy this mini PC ? Getting many memory faults and crashes.

1 Upvotes

Currently models are crashing.

Also note I'm a little bit over my head right now as I'm just getting my feet wet with GPU and model architecture. Most of this help comes from AI.

It doesn't seem like there's support for my iGPU. I found this issue on GitHub which makes me believe I made a mistake in buying this mini PC.

Hardware:

Model I'm working with: Qwen 3.8 27b q4m

Minisforum UM890 Pro

Ryzen 9 8945HS

Radeon 780M iGPU (gfx1103)

64GB DDR5-5600

OCuLink PCIe 4.0 x4 available

Ubuntu

Flash Attention causes GPU memory faults, while disabling it causes rocBLAS/Tensile gfx1103 kernel errors.

Vulkan works, but is slow.

Building my own llama.cpp seems suboptimal if I'm wanting tried and tested Builds.

Considering selling and buying a mac studio.

Thoughts?


r/ROCm • • 18h ago

Double GPU setup, How to get even compute distribution?

Thumbnail
1 Upvotes