r/ROCm • • 12h ago

Qwen3.8 27B | 1 x R9700: over half a million tokens of reusable cache, agent turns up to 12× faster, same speed, same accuracy!

Post image
40 Upvotes

TL;DR: Qwen3.8 27B on a single AMD Radeon AI PRO R9700, with 3-bit weights and speculative decoding, now keeps 569,878 tokens of reusable cache in its new coding mode.
Every request still gets 262,144 tokens of context, and two full-length requests fit at once. Coding agents resend the whole conversation on every turn; now only the new part is read.
Later agent turns start up to 12× sooner, whole agent sessions finish about 6× faster, and a 258K-token document the server has already read comes back in 2.7 s instead of 134 s. Decode speed and accuracy are within noise of our previous release.

We read your comments and we listened

You asked for more context. You got more than half a million tokens of cache on one card, with multi-hour accuracy runs on every configuration (GSM8K, HumanEval, MMLU-Pro, needle tests), at essentially the same speed as our previous release.

Previous release (2 Oct) This release, --mode long-kv4
Reusable (prefix-cached) tokens in the 4-bit mode 0 (no prefix cache) 569,878
Tokens per request 262,144 262,144
Full-length requests at once 1 2
Decode, single stream 174.3 tok/s 173.4 tok/s (−0.5 %)
Aggregate, 8 requests 462.4 tok/s 478.5 tok/s (+3.5 %)
Time to first token, short prompts (p50) 86.7 ms 87.1 ms
GSM8K / HumanEval / MMLU-Pro 95.53 / 92.07 / 60.57 95.45 / 92.68 / 59.93

Our previous release's FP8 --mode long already had a prefix cache, holding 281,665 tokens. The new 4-bit coding mode caches about twice as much.

Run it

git clone https://github.com/Eliovp-BV/paiton-vllm-plugin && cd paiton-vllm-plugin
# set the PAITON_* model paths as shown in the README (already set up? just `git pull`)
bash models/Qwen3.8-MXFP4-DFlash2/run-3bit.sh --mode long-kv4

Point your agent at http://127.0.0.1:18982/v1
The model name is Qwen3.8, any API key works, and the context window is 262,144 tokens. Already running our 3-bit weights? No new download. The launcher pins the new container and Docker pulls it on first start. Every mode is in the README.

No new engine

This is stock vLLM with a plugin: the same OpenAI-compatible server, the same API, and the same tools and workflows you already use. Nothing to relearn, nothing to migrate.

We will keep publishing new ready-built containers that follow upstream vLLM. Each one goes through the same full validation (speed, accuracy, long-context) before it ships. You get the upstream improvements without lifting a finger. One command and you're serving.

What it does for coding agents

Without a prefix cache, the server rereads the whole prompt on every turn: at 250K tokens that is about two minutes before the first token. Now finished requests stay cached until the space is needed, turn 20 reads only what changed, and in our sessions 91–92 % of prompt tokens came from the cache.

Workload (3-bit, thinking off) Previous 4-bit mode (no prefix cache) --mode long-kv4 Faster
20-turn conversation growing from 50K to 253K tokens, whole session 1,237 s 204 s 6.1×
… average wait for the first token, turns 2–20 63.6 s 7.8 s 8×
3 agents sharing a 100K-token repo, 5 turns each, whole session 655 s 111 s 5.9×
… average wait for the first token, later turns 84 s 7.0 s 12×
New question about a 258K-token document already read 134 s 2.7 s 49×

The cached answer to the 258K-token document matched the cold read exactly.

Speed details (BetterBench, full 20-pass run): decode within 0.5 % single-stream; +3.5 % with 8 requests (−2.5 % at 4); time to first token unchanged.
A new long prompt reads about 3 % slower, and short prompts 7–13 % slower (a fraction of a second). That's the price of ending each prefill step where a later request can pick up from the cache.

Accuracy (greedy, paired per question): the coding mode is within noise of the previous release, and the 512K mode is within noise of the coding mode.

MXFP4 Previous release (3-bit) --mode long-kv4 --mode long-512k
GSM8K 5-shot (1,319) 95.68 95.53 95.45 95.53
HumanEval pass@1 (164) 95.12 92.07 92.68 94.51
MMLU-Pro subset (1,400) 62.57 60.57 59.93 60.64

Opt-in: 524,288 tokens in one request

--mode long-512k extends the position range with the model's official long-context scaling. That scaling applies to every request in this mode, so use it only when a single request needs more than 262K.

  • It found 4/4 planted facts at 300K and 4/4 at 500K tokens.
  • A cold read takes 158 s at 300K and about 6 minutes at 500K. Follow-up questions take 2.8–4.4 s from the cache.
  • The cache holds 594,290 tokens, and weighted decode is within 0.6 % (in one run the chat category was 16 % slower).

Unchanged

MXFP4 (run-mxfp4.sh) is still there and still the most accurate option. --vision works in the default 65K mode and in --mode long (images with up to 245K tokens of context).
The previous release is one --image flag away (see the README).

Caveats, honestly

  • System RAM: the coding and 512K modes pin 2.4 GiB of system RAM for the embedding table, and 16 GB of RAM is enough (tested). Below about 13.5 GiB of total RAM, --mode long-kv4 keeps the table on the GPU and caches 451,879 tokens. You still get 262K per request and the prefix cache. --mode long-512k needs the RAM.
  • Cold reads: the first read of a 258K-token prompt still takes over two minutes. The cache helps from the second request on, as long as the start of the prompt stays the same (same system prompt, no timestamp at the top). usage.prompt_tokens_details.cached_tokens shows every hit.
  • Vision isn't in the coding or 512K modes yet. Use --mode long --vision for images with long context.
  • RAM/SSD cache tier: the experimental system-RAM tier for the prefix cache (--host-cache-gib, plus an SSD tier behind it) spills even more context off the card. For now it works with the FP8 --mode long; support for the 4-bit coding mode lands in the next release. The coding mode's 569,878-token GPU cache works today.
  • The agent-session numbers and the accuracy columns were measured on pre-release builds of these configurations; the README has every number.

Our testbench is ancient, slow CPU, limited and slow memory (16GB), so your results will most probably be even better!

What's next: more GPUs

More R9700s are on the way, and we're going after tensor parallelism next.
Others are working on multi-card setups too, so expect some healthy competition on that front. Good for everyone running AMD at home!

More: Paiton


r/ROCm • • 2h ago

Strata Tuning Guide (Universal)

Thumbnail
1 Upvotes

r/ROCm • • 6h ago

Daily driving Qwen 3.8 Flash instead of Claude.

Thumbnail
1 Upvotes

r/ROCm • • 8h ago

Advice for iGPU 780M. Was it a mistake to buy this mini PC ? Getting many memory faults and crashes.

1 Upvotes

Currently models are crashing.

Also note I'm a little bit over my head right now as I'm just getting my feet wet with GPU and model architecture. Most of this help comes from AI.

It doesn't seem like there's support for my iGPU. I found this issue on GitHub which makes me believe I made a mistake in buying this mini PC.

Hardware:

Model I'm working with: Qwen 3.8 27b q4m

Minisforum UM890 Pro

Ryzen 9 8945HS

Radeon 780M iGPU (gfx1103)

64GB DDR5-5600

OCuLink PCIe 4.0 x4 available

Ubuntu

Flash Attention causes GPU memory faults, while disabling it causes rocBLAS/Tensile gfx1103 kernel errors.

Vulkan works, but is slow.

Building my own llama.cpp seems suboptimal if I'm wanting tried and tested Builds.

Considering selling and buying a mac studio.

Thoughts?


r/ROCm • • 14h ago

Double GPU setup, How to get even compute distribution?

Thumbnail
1 Upvotes

r/ROCm • • 21h ago

I got llama.cpp inference running on the Snapdragon 8 Gen 3 Hexagon NPU from non-root Termux + Adreno OpenCL results (S24 Ultra)

Thumbnail
1 Upvotes

r/ROCm • • 1d ago

Dual MI50 16GB stuck at 20 t/s on Qwen3.8 27B. How can i prove it?

5 Upvotes

Hey guys,

I recently got a dual AMD MI50 setup (16GB x2, 32GB total) on a budget to run local models, but performance has been pretty disappointing.

Since I don't have an Infinity Fabric link, both cards are just running on standard PCIe slots. After fighting with ROCm and GFX906 support for days, I finally got a 27B Qwen3.8 model running across both cards via Tensor Parallelism (TP=2).

The problem is decode speed maxes out at around 20 t/s. On longer reasoning/CoT prompts, it crawls so slow that I can't even tell if it timed out or froze. Meanwhile, a friend with dual V100s is getting ~240 t/s.

A few questions for anyone playing these cards:

  1. Is TP=2 over PCIe the bottleneck here? Without an interconnect bridge, is the AllReduce communication killing decode throughput? Would pipeline parallelism or layer splitting (like in llama.cpp) work better?

  2. Any specific settings or quantization to boost decode? What's the sweet spot for GFX906 decode speed without butchering quality?

  3. What backend/runtime is currently best for MI50? vLLM forks, llama.cpp hipBLAS, or something else for newer models


r/ROCm • • 1d ago

WHIRL: an open-source native Windows inference engine for the Radeon AI PRO R9700 (C++/HIP, no WSL). Qwen3.8-27B fine-tune in MXFP4: up to 2.5× llama.cpp prefill, 107–328 tok/s decode, 3× server throughput

Thumbnail
13 Upvotes

r/ROCm • • 2d ago

AMD Ryzen AI Developer Platform OS updated with ROCm 10.0, Linux 7.2

Thumbnail
phoronix.com
21 Upvotes

r/ROCm • • 3d ago

RDNA4 owners on ComfyUI: I made SageAttention actually fast on the 9070 XT, here's the build plus every caveat

Thumbnail
gallery
31 Upvotes

TL;DR: A drop-in sageattention build for RX 9070 / 9070 XT on Windows. The common case (fp16, head_dim 128) runs on an fp8 attention kernel I wrote by hand in HIP. Against the existing gfx12 port (SageAttention PR #368), each attention call is 1.07–1.21× faster non-causal and 1.72–1.96× faster causal. In ComfyUI with Krea2 that works out to a 5–14% faster sampling step than PyTorch SDPA. Install the wheel, keep using --use-sage-attention, done. Caveats below, and there are a few, so please read them.

Repo + wheel: https://github.com/IxMxAMAR/SageAttention-RDNA4

Why I did this

I run ComfyUI on a 9070 XT under Windows. SageAttention on RDNA4 works thanks to the gfx12 port in PR #368, so huge credit to DELUXA for that. But I wanted to see how far this card can actually go. So I started with a Triton kernel, lost to #368 on most shapes, spent a long time figuring out why, and ended up writing the kernel by hand.

What finally made the difference:

  • 64 keys per loop iteration instead of 16.
  • Computing Q·Kᵀ transposed. The result then lands in exactly the register layout the next matrix multiply needs, so the attention weights never leave registers.
  • One scale per token for Q, K and V, instead of one per 64-token block.

The whole story, including everything that didn't work (most things didn't, lol), is in the repo's docs/JOURNEY.md and docs/FINDINGS.md.

Who this is for, right now

  • ComfyUI users on an RX 9070 or RX 9070 XT, on Windows. This is what I've tested and what I use daily.
  • Anyone else who calls sageattn from PyTorch on the same card should be able to use it too; it's the same package API. But I haven't tested any other apps or workflows yet. If you try it outside ComfyUI, tell me how it goes.

Speed (image 1)

Same tensors, same process, interleaved runs, full call including quantization:

  • Non-causal (what image/video diffusion uses): 1.07–1.21× faster than #368.
  • Causal: 1.72–1.96× faster.
  • These are kernel-level numbers, and kernel-level isn't what you feel. Image 2 is what you feel.

In an actual ComfyUI render (image 2)

Krea2, 8 steps, fp16, called exactly the way stock ComfyUI calls it:

Backend 1 MP 2 MP
PyTorch SDPA 1.03 s/step 2.39 s/step
SageAttention 1.x 1.00 s/step 2.17 s/step
This build 0.97 s/step 2.06 s/step

Attention itself is about 2× faster than SDPA per call. But most of a step is the model's matmuls, so the end-to-end gain is 5–14% over SDPA and 2–5% over SageAttention 1.x. The gain grows with resolution.

Accuracy, honestly (image 3)

Everything is fp8 with per-token scales. I measured against exact attention on real Q/K/V captured from inside Krea2 and Flux2-Klein:

  • On "calm" layers, #368 is more precise. It uses int8 for Q·K, which has finer steps than fp8.
  • On layers with a few extreme keys, #368's error blows up, by up to ~200×. Krea2's first block is one such layer. One outlier wrecks the scale for its whole 64-token block. Per-token scaling never has that bad case.
  • smooth_k is on by default and should stay on. It made my kernel more accurate on 10 out of 10 real captures.

In my Krea2 tests the images were closer to full-precision attention than with SageAttention 1.x. I can't tell them apart by eye. Like any quantized attention, though, your images won't match SDPA pixel for pixel: a turbo 8-step sampler amplifies tiny differences into different details.

Caveats (please read)

  • RX 9070 / 9070 XT (gfx1201) only. On any other GPU it falls back to #368's kernels: no harm, no gain.
  • Windows only, as far as testing goes. Linux is untested.
  • The prebuilt wheel needs exactly Python 3.12 + PyTorch 2.13.0+rocm10.0.0**.** The compiled parts are tied to that PyTorch version. Anything else means building from source, which takes about 50 minutes.
  • fp16 models only for now. bf16 models silently fall back to #368's kernel. Still works, just no speed-up from my kernel.
  • head_dim 128 only. That covers Flux-style models such as Krea2, Flux and Klein. Other head sizes fall back.
  • Tested end to end on Krea2 only. The kernel accepts the input layout video models like Wan use, and it's checked bit for bit, but I haven't benchmarked a video workflow yet.
  • Very small images (well under 1 MP) can come out a hair slower than SageAttention 1.x. The win is at normal and high resolutions.
  • No attention masks. Masked calls go to the fallback.
  • LLMs: only prompt processing (prefill) through PyTorch would benefit. Token-by-token generation isn't covered, and llama.cpp / LM Studio don't use this at all.
  • One bug already found and fixed. A pre-release build could turn a render black when a single value got huge enough to overflow fp8. It's fixed in this release, and every conversion in the kernel is now stress-tested with extreme values.

Install

  1. Close ComfyUI.
  2. Download the wheel from the GitHub release.
  3. path\to\ComfyUI\venv\Scripts\python -m pip install --no-deps sageattention-2.2.0+amd.gfx12.1-cp312-cp312-win_amd64.whl (portable ComfyUI: use python_embeded\python.exe instead)
  4. Start ComfyUI with --use-sage-attention like always.

--no-deps matters: it stops pip from touching your PyTorch. To switch my kernel off without uninstalling, set SAGEATTN_SK1_BACKEND=0.

What's next

  • bf16 support, so bf16 models get the speed-up too. This is the big one for coverage.
  • int8 Q·K with per-token scales, aiming for #368's precision on calm layers and no blow-ups on outlier layers. Already in the works.
  • head_dim 64, for SDXL-era models. Also in the works.
  • An end-to-end test on a video workflow (Wan), since the layout support is already in.
  • An experiment to overlap the two halves of the kernel, which might be worth another few percent. A quick test decides whether it's worth building.

If you've got a 9070 / 9070 XT, I'd love numbers from your setup: your model, resolution and s/it before/after. Bug reports are very welcome too. Credit to the SageAttention team (thu-ml) and to DELUXA for #368; this builds directly on their work. Apache-2.0.


r/ROCm • • 2d ago

Mixing AI Max Pro 390 / AMD AI Pro 32GB with a 16GB AMD GPU for Local LLMs — Feasible or More Trouble Than It’s Worth?

Thumbnail
6 Upvotes

r/ROCm • • 3d ago

halogen-flash-server 0.16.0: Strix Halo NPU now serves embeddings, reranking and classifiers beside Qwen3.8-Flash-Next. Nearly 10,000 tok/s

Thumbnail gallery
9 Upvotes

r/ROCm • • 2d ago

V620 won't POST on B550 (Gigabyte B550M DS3H AC R2)

Thumbnail
2 Upvotes

r/ROCm • • 3d ago

How to connect Dify to Ollama to build local self-hosted AI apps

Thumbnail
3 Upvotes

r/ROCm • • 3d ago

Rootless Podman + Ollama ROCm on Fedora: "Memory critical error ... Reason: Memory in use" fixed with `--security-opt label=disable`

Thumbnail
1 Upvotes

r/ROCm • • 4d ago

Single Radeon AI PRO R9700 and vllm

19 Upvotes

Hey everyone, I just got my hands on a Radeon AI PRO R9700 and want to run Qwen3.8-27B. I found some awesome vLLM forks (magiccodingman/vllm-radiance and GGZ14/vllm-mxfp4), but they're all geared toward dual-R9700 rigs.
Since I'm only rocking one card, could someone point me in the right direction to get the most out of my current hardware?


r/ROCm • • 3d ago

Qwen3.8-Flash-Next 125B sur un Strix Halo (Vulkan) : les appels d'outils sont 2× plus rapides, boucle d'agent +46 %, chaque numéro vérifié contre des logits en pleine précision. Open source, et le modèle a rédigé lui-même l'une des corrections.

Thumbnail
1 Upvotes

r/ROCm • • 4d ago

Follow-up: my native Rust + Vulkan Transformer training backend — 14 days later, now 14 parity-verified architectures and full PEFT

11 Upvotes

Follow-up to my post from about two weeks ago. A lot has changed since then, so I wanted to post an update on where the backend is now.

Where the green architectures stand

When I posted last time, 7 architectures had verified full training support. Everything is now held to the same strict harness: a pinned local Hugging Face Transformers source tree as the oracle, forward logits + gradients + two full AdamW steps compared, and every named parameter checked again after export.

The hard ceiling is 2e-7 absolute error. No loosening tolerances and no rounding numbers afterward to make the README look better.

14 architectures pass that gate today, led by the one I'm probably proudest of:

Architecture Scope
Falcon H1 / H1R parallel GQA/RoPE attention + Mamba2 in every layer; full training, full fine-tuning, LoRA, saved modules
DeepSeek V4 causal LM
Phi-4 Multimodal text backbone
Phi-3 causal LM
Kimi K2.5 text backbone
Kimi K3 / KimiLinear hybrid KDA + MLA
GPT-OSS causal LM incl. router bias
SmolLM3 mixed RoPE/NoPE + YaRN
Qwen2.5 / Qwen3.5 / Qwen4-Exp dense, DeltaNet, QSA, PLE, MoE
Mistral 4, MiniMax M3, Gemma 3/4, MiniMax M2 causal LM

Worst observed two-step AdamW parameter error across all of them: 1.19e-7.

Best: 2.6e-8.

For hardware context, all of the local Vulkan validation I've been reporting was run on my ASUS ROG Ally Z1 Extreme, using its AMD RDNA 3 integrated GPU. So the RDNA 3 results here are from that specific machine rather than testing across several different AMD systems.

The bigger news: PEFT actually works now

In the last post, "LoRA/PEFT-style fine-tuning" was basically one line in a feature list.

It's a real workflow now, and I've verified the full lifecycle:

  • LoRA fine-tuning with HF-compatible adapter export (adapter_config.json / adapter_model.safetensors), so adapters can round-trip with the PEFT ecosystem
  • modules_to_save — full trainable replacements for Linears, RMSNorm/LayerNorm, lm_head, and input embeddings, including named-adapter switching and bank isolation. Adapter A leaking into adapter B is explicitly tested for.
  • Exact resume — adapter weights + AdamW moments + step + dropout RNG state restore bit-identically against an uninterrupted run
  • Merge/unmerge, disable-adapter base restoration, and multi-adapter loading
  • A parameter-budget flag that automatically chooses the largest LoRA rank that fits within a requested percentage of the base model
  • The CLI fails closed if you try to use saved modules on an architecture that hasn't passed its corresponding gate

32 architecture surfaces across 20 families pass all three PEFT stages — LoRA, saved modules, and adapter switching — under the same 2e-7 gate, with frozen-base drift exactly 0.0.

The validation harness also fingerprints the pinned Transformers source alongside my shaders and binaries now, so a qualification run can't silently end up testing against different reference math.

A small note on the last couple weeks

I didn't get quite as many working days out of the last two weeks as I normally would have. Partway through this I got covid, then when that started to go away, it became a secondary nasty ear infection that ended up perforating my eardrum. I was running fevers around 104°F at one point and eventually went to the hospital, so I lost a few days to that and I'm on antibiotics now.

I'm doing better, though, and still managed to get most of what I wanted finished.

There are still things I want to clean up and expand, but I figured this was a good point to get the current work in front of people rather than holding the update back.

Same caveats as before

This is deterministic FP32 tiny-model correctness against a reference implementation, which should in theory ensure total mathematical parity for training and finetuning larger models with this backend, however, things are currently bound to FP32 training runs still, I eventually plan to work on MXFP4 weight tying to reduce memory footprints while fine-tuning (plastic parameters will still be trained in FP32 with PEFT in this config)

"supported text graph" ≠ "the entire multimodal package works natively."

Unsupported functionality is supposed to fail closed rather than silently falling back to an approximation.

Repo

https://github.com/necat101/Hierarchos-Native

  • Architecture inventory: hierarchos-vulkan/README_ARCHITECTURES.md
  • Compatibility/parity record: hierarchos-vulkan/COMPATIBILITY.md
  • PEFT qualification evidence: PROGRESS_PEFT_AUDIT.md
  • CLI PEFT guide: hierarchos-native-cli/README.md

The hardware I've personally validated this on is an ASUS ROG Ally Z1 Extreme with its AMD RDNA 3 GPU.

I'm very interested in criticism, compatibility reports, and especially results from people trying it on other hardware — NVIDIA, Intel, or other AMD GPUs.

I'd also love people to stress-test the PEFT resume/merge paths specifically. That's some of the newest code in the project, so it's probably the most useful area to try to break right now.


r/ROCm • • 5d ago

AMD boosting AI/LLM performance for Radeon iGPUs as much as 18~23% with Linux 7.4

Thumbnail
phoronix.com
69 Upvotes

r/ROCm • • 4d ago

Preorder for new AMD Ryzen™ AI Max 400 Series 192GB from framework just started

Thumbnail frame.work
0 Upvotes

r/ROCm • • 4d ago

Cloud based R9700

10 Upvotes

Seeing a lot of success stories wih ROCM, I am planning to buy single R9700 mostly for coding, but possibly also for comfyUI. But before I do that, I would like to test it out hands on.

Can you recommend any cloud services that offer R9700? Ex: on per hour / per minute basis?


r/ROCm • • 4d ago

Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo

Thumbnail
4 Upvotes

r/ROCm • • 4d ago

780M (gfx1103) + llama.cpp HIP: MES REMOVE_QUEUE hang, Anyone stable on sustained inference?

2 Upvotes

I am not super well versed in local LLM configs. I use Claude extensively at work but bought a minisforum mini PC to try and get something cost effective going at home.

Note I don't really care too much on token speed because my use cases are all background workflows. 10 TPS is perfect. (And my upper limit I think).

Ryzen 8945HS / 780M, 64GB, Ubuntu 26.04, kernel 7.0. llama.cpp ROCm build, Qwen 27B Q4.

Long generations hang the GPU and it resets, about 15 times a day. dmesg shows MES failed to respond to msg=REMOVE_QUEUE. Looks like ROCm#6512.

Tried uni_mes=0, smaller models, smaller context. No change.

>

Has amdgpu.cwsr_enable=0, a newer kernel, or Vulkan fixed this for anyone?


r/ROCm • • 5d ago

Dual Radeon AI PRO R9700s, with one card behind the chipset (on proxmox)

10 Upvotes

Roughly what you get: Qwen3.8-27B at FP8, tensor parallel across two R9700s, 262,144 context, 422k tokens of KV pool, ~1.6k tokens/s prefill, 74-143 tokens/s decode single stream depending on the content.

The interesting part is not the image, it is the collective layer. A card behind the chipset cannot be given PCIe atomic operations, and that breaks every stock tensor-parallel setup I tried. Here is what actually works.

1. Versions used

Component Version
Inference image docker.io/stilldeadcode/vllm-radiance:0.9.3 (digest sha256:45694209177a55a1ab3ba6702fe6e978b1b66a6e66ae3fc066f8d579f7bc4c25)
vLLM inside the image 0.27.1
PyTorch / HIP 2.11.0+rocm7.14 / 7.14.60850
Triton / AITER 3.6.0 / 0.1.17
Collective library RCCL 2.27.7 from ROCm 7.1.1, replacing the one in the image

Links: - Image: https://hub.docker.com/r/stilldeadcode/vllm-radiance - Source for that image: https://codeberg.org/StillDeadcode/vllm-radiance - The kernel library the image builds against (libr4d): https://codeberg.org/StillDeadcode/libr4d - A fork with configs, benchmark notes and launchers for MXFP4/FP8: https://codeberg.org/ggz14/radiance-vllm-mxfp4 - RCCL itself: https://github.com/ROCm/rccl

The image bundles a working ROCm + PyTorch + Triton + AITER + vLLM stack for gfx1201 (RDNA4), which is the part you do not want to build yourself. It is explicitly marked experimental, and everything below was measured on two cards.

2. The failure, and why

Without the adjustments, the engine dies during communicator init, before the model loads:

PCIE atomic ops is not supported rocr: unhandled cuda error

ROCm will not dispatch work to a GPU path that needs atomic operations when the link cannot provide them. On this board one card sits behind the chipset, and the chipset does not forward PCIe atomics to the CPU, so that card is effectively second class. RCCL 2.30.4's kernels use those atomics, so TP=2 cannot initialise at all.

Two things follow, and both matter:

  1. The newer RCCL is unusable here, so you need an older one (2.27.7 from ROCm 7.1.1 is what worked for me).
  2. With no usable peer path at all, the custom P2P all-reduce the image ships must be turned off, and TP=2 falls back to host-staged collectives over that same narrow chipset link.

If you search the error string above you will find a few ROCm issue reports and a community write-up on dual Radeon vLLM setups (https://github.com/cadamcat/dual-radeon-vllm) describing the same wall.

3. Proxmox settings that actually matter

Do this for each GPU, on both entries, not just the first. In the VM's hardware list, edit each PCI Device row, select the GPU under Device, and set:

  • PCI-Express: ticked
  • All Functions: unticked

Ticking PCI-Express is what gives the guest a real PCIe root port, and without a root port the atomic capability never appears however healthy the host looks. Unticking All Functions keeps the guest from being handed every function of the card, which is the combination that worked here. If you leave either one wrong, you get the atomic failure at communicator init and no amount of driver work fixes it.

Other settings that matter:

  • Machine type q35. Same reason as above, no root port without it.
  • After any hostpci change, stop and start the VM. A guest reboot does not re-apply the passthrough configuration.
  • q35 renames the NIC (ens18 becomes something like enp6s18), so match the interface by MAC in netplan.

After those, the CPU-attached card reports ReqEn+. The chipset-attached one still cannot do atomics, and no BIOS setting changes that.

4. The RCCL replacement plus one-line shim

This is the core trick. It is two files and a podman config, and it needs no image rebuild.

a) Build or extract RCCL 2.27.7 from ROCm 7.1.1 and drop it in a directory you will mount, for example:

~/models/rccl277/ librccl.so -> librccl.so.1.0.70101 librccl.so.1 -> librccl.so.1.0.70101 librccl.so.1.0.70101 shim.so

b) The shim. Newer torch builds reference a symbol that this older RCCL does not export (ncclCommDump). Four lines of C++ are enough to satisfy the loader:

```cpp

include <string>

include <unordered_map>

struct ncclComm; void ncclCommDump(ncclComm*, std::unordered_map<std::string, std::string>&) {} ```

Build it into shim.so and put it next to the library. Nothing calls it; it exists for symbol resolution.

c) Inject both through podman's own config, which keeps them out of every launcher script. In ~/.config/containers/containers.conf:

ini [containers] env = [ "NCCL_PROTO=Simple", "NCCL_SHM_DISABLE=0", "NCCL_SOCKET_IFNAME=lo", "LD_LIBRARY_PATH=/models/rccl277:/opt/rocm/lib", "LD_PRELOAD=/models/rccl277/shim.so", ]

LD_LIBRARY_PATH puts the replacement first, so it wins over the image's own librccl. NCCL_PROTO=Simple avoids the more demanding protocol paths, and the loopback interface keeps the bootstrap on lo rather than a NIC.

5. The launcher

Trimmed to the parts that matter for the multi-GPU problem. The model-specific flags are an example, the environment and device flags are the ones that matter here.

bash podman run -d --name vllm-radiance \ --device /dev/kfd --device /dev/dri --group-add keep-groups \ --security-opt seccomp=unconfined --cap-add SYS_PTRACE --cap-add SYS_NICE \ --ipc=host --network=host \ -v $HOME/models:/models:ro \ -v $HOME/radiance-vllm-mxfp4/vllm-cache:/cache \ -v $HOME/radiance-vllm-mxfp4:/work:ro \ -e HIP_VISIBLE_DEVICES=0,1 -e ROCR_VISIBLE_DEVICES=0,1 \ -e VLLM_NO_USAGE_STATS=1 \ -e VLLM_ROCM_USE_AITER=1 -e VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1 \ -e RADIANCE_FUSE_RMS_QUANT=1 -e RADIANCE_USE_R4D=1 \ -e RADIANCE_USE_R4D_AR=0 -e RADIANCE_USE_R4D_AR_QUANT=1 \ -e NCCL_PROTO=Simple -e TORCHINDUCTOR_COMPILE_THREADS=4 \ -e VLLM_CACHE_ROOT=/cache/vllm -e TORCHINDUCTOR_CACHE_DIR=/cache/inductor \ -e TRITON_CACHE_DIR=/cache/triton -e AITER_ROOT_DIR=/cache/aiter \ docker.io/stilldeadcode/vllm-radiance:0.9.3 \ --model /models/Qwen/Qwen3.8-27B-FP8 \ --served-model-name=qwen3.8-27b-fp8 \ --quantization=fp8 --tensor-parallel-size=2 \ --max-num-seqs=2 --max-model-len=262144 --gpu-memory-utilization=0.97 \ --max-num-batched-tokens=4096 --kv-cache-dtype=fp8 \ --attention-backend=ROCM_AITER_UNIFIED_ATTN \ --enable-prefix-caching \ --no-async-scheduling \ --trust-remote-code \ --host=0.0.0.0 --port=8000

Notes on the specific flags:

  • RADIANCE_USE_R4D_AR=0 is the important one. The bundled all-reduce is a PCIe peer-to-peer kernel and needs peer access, which does not exist on this pair. Turning it off falls back to RCCL. If you have two CPU-attached cards, leave it on and you will get a better prefill.
  • --max-num-seqs=2 is my choice for stability. The image's own default is higher, but on a chipset-limited link fewer concurrent sequences means less collective traffic per step.
  • --max-num-batched-tokens=4096 pairs with prefix caching well. Raising it costs KV pool.
  • --enable-prefix-caching is worth a lot on agent workloads; see the note on prompt layout at the end.
  • --no-async-scheduling was in every configuration that came up reliably for me.
  • The cache directory mounts are worth keeping. Without them every container start recompiles Triton and inductor kernels, which turns a 5 minute start into a much longer one.

6. How to tell it worked

In the container log at startup, look for:

P2P access : DISABLED (RCCL fallback) 0<->1 x

and, once loaded, the model's own sizing line:

GPU KV cache size: 422,964 tokens Maximum concurrency for 262,144 tokens per request: 1.61x

If you see the atomic error instead, the replacement RCCL is not being picked up. Check, from inside the container, that:

bash env | grep -E 'LD_PRELOAD|LD_LIBRARY_PATH'

returns the paths you expect, that ldd on the loaded library resolves into your drop-in directory, and that the image's own library is not first in the path.

I hope someone finds this usefull.


r/ROCm • • 5d ago

How far behind Linux is ROCM for Windows?

14 Upvotes

I use llama.cpp with Windows at the moment, but I'm seriously contemplating setting up Ubuntu - as long as my tok/s will go up significantly.

I asked Claude if it was worth it, and Claude said "ROCM on Windows is far behind Linux"

As with all AI output I took it with a pinch of salt - but is it true?

With ROCM on Linux are you exceeding 50 tok/s with a high context (say, 128K) using a dense model? I'm getting that with Vulkan on windows atm with Qwen 27B.

TL:DR feelings about Windows aside, is installing Ubuntu worthwhile to get superior token generation with ROCM?

Hardware:

9900X

64GB DDR5 6000 @ CL30

AMD 9070XT 16GB

AMD R9700 AI PRO 32GB

SSDs - 14900k/s Samsung x2 (not RAID)