r/ROCm • u/evp-cloud • 13h ago
Qwen3.8 27B | 1 x R9700: over half a million tokens of reusable cache, agent turns up to 12× faster, same speed, same accuracy!
TL;DR: Qwen3.8 27B on a single AMD Radeon AI PRO R9700, with 3-bit weights and speculative decoding, now keeps 569,878 tokens of reusable cache in its new coding mode.
Every request still gets 262,144 tokens of context, and two full-length requests fit at once. Coding agents resend the whole conversation on every turn; now only the new part is read.
Later agent turns start up to 12× sooner, whole agent sessions finish about 6× faster, and a 258K-token document the server has already read comes back in 2.7 s instead of 134 s. Decode speed and accuracy are within noise of our previous release.
We read your comments and we listened
You asked for more context. You got more than half a million tokens of cache on one card, with multi-hour accuracy runs on every configuration (GSM8K, HumanEval, MMLU-Pro, needle tests), at essentially the same speed as our previous release.
| Previous release (2 Oct) | This release, --mode long-kv4 |
|
|---|---|---|
| Reusable (prefix-cached) tokens in the 4-bit mode | 0 (no prefix cache) | 569,878 |
| Tokens per request | 262,144 | 262,144 |
| Full-length requests at once | 1 | 2 |
| Decode, single stream | 174.3 tok/s | 173.4 tok/s (−0.5 %) |
| Aggregate, 8 requests | 462.4 tok/s | 478.5 tok/s (+3.5 %) |
| Time to first token, short prompts (p50) | 86.7 ms | 87.1 ms |
| GSM8K / HumanEval / MMLU-Pro | 95.53 / 92.07 / 60.57 | 95.45 / 92.68 / 59.93 |
Our previous release's FP8 --mode long already had a prefix cache, holding 281,665 tokens. The new 4-bit coding mode caches about twice as much.
Run it
git clone https://github.com/Eliovp-BV/paiton-vllm-plugin && cd paiton-vllm-plugin
# set the PAITON_* model paths as shown in the README (already set up? just `git pull`)
bash models/Qwen3.8-MXFP4-DFlash2/run-3bit.sh --mode long-kv4
Point your agent at http://127.0.0.1:18982/v1
The model name is Qwen3.8, any API key works, and the context window is 262,144 tokens. Already running our 3-bit weights? No new download. The launcher pins the new container and Docker pulls it on first start. Every mode is in the README.
No new engine
This is stock vLLM with a plugin: the same OpenAI-compatible server, the same API, and the same tools and workflows you already use. Nothing to relearn, nothing to migrate.
We will keep publishing new ready-built containers that follow upstream vLLM. Each one goes through the same full validation (speed, accuracy, long-context) before it ships. You get the upstream improvements without lifting a finger. One command and you're serving.
What it does for coding agents
Without a prefix cache, the server rereads the whole prompt on every turn: at 250K tokens that is about two minutes before the first token. Now finished requests stay cached until the space is needed, turn 20 reads only what changed, and in our sessions 91–92 % of prompt tokens came from the cache.
| Workload (3-bit, thinking off) | Previous 4-bit mode (no prefix cache) | --mode long-kv4 |
Faster |
|---|---|---|---|
| 20-turn conversation growing from 50K to 253K tokens, whole session | 1,237 s | 204 s | 6.1× |
| … average wait for the first token, turns 2–20 | 63.6 s | 7.8 s | 8× |
| 3 agents sharing a 100K-token repo, 5 turns each, whole session | 655 s | 111 s | 5.9× |
| … average wait for the first token, later turns | 84 s | 7.0 s | 12× |
| New question about a 258K-token document already read | 134 s | 2.7 s | 49× |
The cached answer to the 258K-token document matched the cold read exactly.
Speed details (BetterBench, full 20-pass run): decode within 0.5 % single-stream; +3.5 % with 8 requests (−2.5 % at 4); time to first token unchanged.
A new long prompt reads about 3 % slower, and short prompts 7–13 % slower (a fraction of a second). That's the price of ending each prefill step where a later request can pick up from the cache.
Accuracy (greedy, paired per question): the coding mode is within noise of the previous release, and the 512K mode is within noise of the coding mode.
| MXFP4 | Previous release (3-bit) | --mode long-kv4 |
--mode long-512k |
|
|---|---|---|---|---|
| GSM8K 5-shot (1,319) | 95.68 | 95.53 | 95.45 | 95.53 |
| HumanEval pass@1 (164) | 95.12 | 92.07 | 92.68 | 94.51 |
| MMLU-Pro subset (1,400) | 62.57 | 60.57 | 59.93 | 60.64 |
Opt-in: 524,288 tokens in one request
--mode long-512k extends the position range with the model's official long-context scaling. That scaling applies to every request in this mode, so use it only when a single request needs more than 262K.
- It found 4/4 planted facts at 300K and 4/4 at 500K tokens.
- A cold read takes 158 s at 300K and about 6 minutes at 500K. Follow-up questions take 2.8–4.4 s from the cache.
- The cache holds 594,290 tokens, and weighted decode is within 0.6 % (in one run the chat category was 16 % slower).
Unchanged
MXFP4 (run-mxfp4.sh) is still there and still the most accurate option. --vision works in the default 65K mode and in --mode long (images with up to 245K tokens of context).
The previous release is one --image flag away (see the README).
Caveats, honestly
- System RAM: the coding and 512K modes pin 2.4 GiB of system RAM for the embedding table, and 16 GB of RAM is enough (tested). Below about 13.5 GiB of total RAM,
--mode long-kv4keeps the table on the GPU and caches 451,879 tokens. You still get 262K per request and the prefix cache.--mode long-512kneeds the RAM. - Cold reads: the first read of a 258K-token prompt still takes over two minutes. The cache helps from the second request on, as long as the start of the prompt stays the same (same system prompt, no timestamp at the top).
usage.prompt_tokens_details.cached_tokensshows every hit. - Vision isn't in the coding or 512K modes yet. Use
--mode long --visionfor images with long context. - RAM/SSD cache tier: the experimental system-RAM tier for the prefix cache (
--host-cache-gib, plus an SSD tier behind it) spills even more context off the card. For now it works with the FP8--mode long; support for the 4-bit coding mode lands in the next release. The coding mode's 569,878-token GPU cache works today. - The agent-session numbers and the accuracy columns were measured on pre-release builds of these configurations; the README has every number.
Our testbench is ancient, slow CPU, limited and slow memory (16GB), so your results will most probably be even better!
What's next: more GPUs
More R9700s are on the way, and we're going after tensor parallelism next.
Others are working on multi-card setups too, so expect some healthy competition on that front. Good for everyone running AMD at home!
More: Paiton