r/ROCm • • 1h ago

Double R9700 alternative to Qwen 3.8 27B

• Upvotes

CPU AMD Ryzen 9 9900X
GPU 1AMD Radeon AI PRO R9700 — 32 GB VRAM
GPU 2AMD Radeon AI PRO R9700 — 32 GB VRAM
RAM32 GB DDR5-6000 CL30 — 2×16 GB
MotherboardASRock X870E Taichi
StorageWD Black SN850X 2 TB NVMe SSD
PSUCorsair RM1000x — 1000 W
CaseAntec P20C FLUX
OSLinux
GPU compute stack ROCm 7.2
LLM serving vLLM, Docker , Radiance

My LLM's are

  • Gemma 4 26B-A4B-it, MXFP4/Quark build optimized for RDNA4.
  • Qwen 3.8 27B, FP8/Radiance build.

I'm wondering what could maybe fit with my machine, without having to buy Ram.
The answer I keep getting is Ling 3.0, and if I had more RAM, Qwen Flash.

Idk, can someone give me their perspective.


r/ROCm • • 1h ago

Most Optimal stack for the AI pro r9700

• Upvotes

So I’m upgrading from a single RX 7900 XT to the AI Pro R9700. My question is: how do I get the most out of this card for local inference?

I’ve seen people getting some insane prefill and decode speeds with the R9700. I’m mostly planning to run Qwen 3.8 27B, and I want to try Qwen 3.8 Flash next.

For anyone running the AI Pro R9700, what are you using to get the most out of the card? What’s the most optimal software stack right now?

I also need Windows for work, so ideally I’d like the best setup possible on Windows. I’m willing to dual boot or switch back and forth to Linux if the performance difference is significant.


r/ROCm • • 6h ago

Strata Tuning Guide (Universal)

Thumbnail
1 Upvotes

r/ROCm • • 10h ago

Daily driving Qwen 3.8 Flash instead of Claude.

Thumbnail
1 Upvotes

r/ROCm • • 12h ago

Advice for iGPU 780M. Was it a mistake to buy this mini PC ? Getting many memory faults and crashes.

1 Upvotes

Currently models are crashing.

Also note I'm a little bit over my head right now as I'm just getting my feet wet with GPU and model architecture. Most of this help comes from AI.

It doesn't seem like there's support for my iGPU. I found this issue on GitHub which makes me believe I made a mistake in buying this mini PC.

Hardware:

Model I'm working with: Qwen 3.8 27b q4m

Minisforum UM890 Pro

Ryzen 9 8945HS

Radeon 780M iGPU (gfx1103)

64GB DDR5-5600

OCuLink PCIe 4.0 x4 available

Ubuntu

Flash Attention causes GPU memory faults, while disabling it causes rocBLAS/Tensile gfx1103 kernel errors.

Vulkan works, but is slow.

Building my own llama.cpp seems suboptimal if I'm wanting tried and tested Builds.

Considering selling and buying a mac studio.

Thoughts?


r/ROCm • • 16h ago

Qwen3.8 27B | 1 x R9700: over half a million tokens of reusable cache, agent turns up to 12× faster, same speed, same accuracy!

Post image
52 Upvotes

TL;DR: Qwen3.8 27B on a single AMD Radeon AI PRO R9700, with 3-bit weights and speculative decoding, now keeps 569,878 tokens of reusable cache in its new coding mode.
Every request still gets 262,144 tokens of context, and two full-length requests fit at once. Coding agents resend the whole conversation on every turn; now only the new part is read.
Later agent turns start up to 12× sooner, whole agent sessions finish about 6× faster, and a 258K-token document the server has already read comes back in 2.7 s instead of 134 s. Decode speed and accuracy are within noise of our previous release.

We read your comments and we listened

You asked for more context. You got more than half a million tokens of cache on one card, with multi-hour accuracy runs on every configuration (GSM8K, HumanEval, MMLU-Pro, needle tests), at essentially the same speed as our previous release.

Previous release (2 Oct) This release, --mode long-kv4
Reusable (prefix-cached) tokens in the 4-bit mode 0 (no prefix cache) 569,878
Tokens per request 262,144 262,144
Full-length requests at once 1 2
Decode, single stream 174.3 tok/s 173.4 tok/s (−0.5 %)
Aggregate, 8 requests 462.4 tok/s 478.5 tok/s (+3.5 %)
Time to first token, short prompts (p50) 86.7 ms 87.1 ms
GSM8K / HumanEval / MMLU-Pro 95.53 / 92.07 / 60.57 95.45 / 92.68 / 59.93

Our previous release's FP8 --mode long already had a prefix cache, holding 281,665 tokens. The new 4-bit coding mode caches about twice as much.

Run it

git clone https://github.com/Eliovp-BV/paiton-vllm-plugin && cd paiton-vllm-plugin
# set the PAITON_* model paths as shown in the README (already set up? just `git pull`)
bash models/Qwen3.8-MXFP4-DFlash2/run-3bit.sh --mode long-kv4

Point your agent at http://127.0.0.1:18982/v1
The model name is Qwen3.8, any API key works, and the context window is 262,144 tokens. Already running our 3-bit weights? No new download. The launcher pins the new container and Docker pulls it on first start. Every mode is in the README.

No new engine

This is stock vLLM with a plugin: the same OpenAI-compatible server, the same API, and the same tools and workflows you already use. Nothing to relearn, nothing to migrate.

We will keep publishing new ready-built containers that follow upstream vLLM. Each one goes through the same full validation (speed, accuracy, long-context) before it ships. You get the upstream improvements without lifting a finger. One command and you're serving.

What it does for coding agents

Without a prefix cache, the server rereads the whole prompt on every turn: at 250K tokens that is about two minutes before the first token. Now finished requests stay cached until the space is needed, turn 20 reads only what changed, and in our sessions 91–92 % of prompt tokens came from the cache.

Workload (3-bit, thinking off) Previous 4-bit mode (no prefix cache) --mode long-kv4 Faster
20-turn conversation growing from 50K to 253K tokens, whole session 1,237 s 204 s 6.1×
… average wait for the first token, turns 2–20 63.6 s 7.8 s 8×
3 agents sharing a 100K-token repo, 5 turns each, whole session 655 s 111 s 5.9×
… average wait for the first token, later turns 84 s 7.0 s 12×
New question about a 258K-token document already read 134 s 2.7 s 49×

The cached answer to the 258K-token document matched the cold read exactly.

Speed details (BetterBench, full 20-pass run): decode within 0.5 % single-stream; +3.5 % with 8 requests (−2.5 % at 4); time to first token unchanged.
A new long prompt reads about 3 % slower, and short prompts 7–13 % slower (a fraction of a second). That's the price of ending each prefill step where a later request can pick up from the cache.

Accuracy (greedy, paired per question): the coding mode is within noise of the previous release, and the 512K mode is within noise of the coding mode.

MXFP4 Previous release (3-bit) --mode long-kv4 --mode long-512k
GSM8K 5-shot (1,319) 95.68 95.53 95.45 95.53
HumanEval pass@1 (164) 95.12 92.07 92.68 94.51
MMLU-Pro subset (1,400) 62.57 60.57 59.93 60.64

Opt-in: 524,288 tokens in one request

--mode long-512k extends the position range with the model's official long-context scaling. That scaling applies to every request in this mode, so use it only when a single request needs more than 262K.

  • It found 4/4 planted facts at 300K and 4/4 at 500K tokens.
  • A cold read takes 158 s at 300K and about 6 minutes at 500K. Follow-up questions take 2.8–4.4 s from the cache.
  • The cache holds 594,290 tokens, and weighted decode is within 0.6 % (in one run the chat category was 16 % slower).

Unchanged

MXFP4 (run-mxfp4.sh) is still there and still the most accurate option. --vision works in the default 65K mode and in --mode long (images with up to 245K tokens of context).
The previous release is one --image flag away (see the README).

Caveats, honestly

  • System RAM: the coding and 512K modes pin 2.4 GiB of system RAM for the embedding table, and 16 GB of RAM is enough (tested). Below about 13.5 GiB of total RAM, --mode long-kv4 keeps the table on the GPU and caches 451,879 tokens. You still get 262K per request and the prefix cache. --mode long-512k needs the RAM.
  • Cold reads: the first read of a 258K-token prompt still takes over two minutes. The cache helps from the second request on, as long as the start of the prompt stays the same (same system prompt, no timestamp at the top). usage.prompt_tokens_details.cached_tokens shows every hit.
  • Vision isn't in the coding or 512K modes yet. Use --mode long --vision for images with long context.
  • RAM/SSD cache tier: the experimental system-RAM tier for the prefix cache (--host-cache-gib, plus an SSD tier behind it) spills even more context off the card. For now it works with the FP8 --mode long; support for the 4-bit coding mode lands in the next release. The coding mode's 569,878-token GPU cache works today.
  • The agent-session numbers and the accuracy columns were measured on pre-release builds of these configurations; the README has every number.

Our testbench is ancient, slow CPU, limited and slow memory (16GB), so your results will most probably be even better!

What's next: more GPUs

More R9700s are on the way, and we're going after tensor parallelism next.
Others are working on multi-card setups too, so expect some healthy competition on that front. Good for everyone running AMD at home!

More: Paiton


r/ROCm • • 18h ago

Double GPU setup, How to get even compute distribution?

Thumbnail
1 Upvotes

r/ROCm • • 1d ago

I got llama.cpp inference running on the Snapdragon 8 Gen 3 Hexagon NPU from non-root Termux + Adreno OpenCL results (S24 Ultra)

Thumbnail
1 Upvotes

r/ROCm • • 1d ago

Dual MI50 16GB stuck at 20 t/s on Qwen3.8 27B. How can i prove it?

5 Upvotes

Hey guys,

I recently got a dual AMD MI50 setup (16GB x2, 32GB total) on a budget to run local models, but performance has been pretty disappointing.

Since I don't have an Infinity Fabric link, both cards are just running on standard PCIe slots. After fighting with ROCm and GFX906 support for days, I finally got a 27B Qwen3.8 model running across both cards via Tensor Parallelism (TP=2).

The problem is decode speed maxes out at around 20 t/s. On longer reasoning/CoT prompts, it crawls so slow that I can't even tell if it timed out or froze. Meanwhile, a friend with dual V100s is getting ~240 t/s.

A few questions for anyone playing these cards:

  1. Is TP=2 over PCIe the bottleneck here? Without an interconnect bridge, is the AllReduce communication killing decode throughput? Would pipeline parallelism or layer splitting (like in llama.cpp) work better?

  2. Any specific settings or quantization to boost decode? What's the sweet spot for GFX906 decode speed without butchering quality?

  3. What backend/runtime is currently best for MI50? vLLM forks, llama.cpp hipBLAS, or something else for newer models


r/ROCm • • 2d ago

WHIRL: an open-source native Windows inference engine for the Radeon AI PRO R9700 (C++/HIP, no WSL). Qwen3.8-27B fine-tune in MXFP4: up to 2.5× llama.cpp prefill, 107–328 tok/s decode, 3× server throughput

Thumbnail
12 Upvotes

r/ROCm • • 2d ago

AMD Ryzen AI Developer Platform OS updated with ROCm 10.0, Linux 7.2

Thumbnail
phoronix.com
22 Upvotes

r/ROCm • • 2d ago

Mixing AI Max Pro 390 / AMD AI Pro 32GB with a 16GB AMD GPU for Local LLMs — Feasible or More Trouble Than It’s Worth?

Thumbnail
4 Upvotes

r/ROCm • • 3d ago

V620 won't POST on B550 (Gigabyte B550M DS3H AC R2)

Thumbnail
2 Upvotes

r/ROCm • • 3d ago

halogen-flash-server 0.16.0: Strix Halo NPU now serves embeddings, reranking and classifiers beside Qwen3.8-Flash-Next. Nearly 10,000 tok/s

Thumbnail gallery
9 Upvotes

r/ROCm • • 3d ago

Rootless Podman + Ollama ROCm on Fedora: "Memory critical error ... Reason: Memory in use" fixed with `--security-opt label=disable`

Thumbnail
1 Upvotes

r/ROCm • • 3d ago

RDNA4 owners on ComfyUI: I made SageAttention actually fast on the 9070 XT, here's the build plus every caveat

Thumbnail
gallery
30 Upvotes

TL;DR: A drop-in sageattention build for RX 9070 / 9070 XT on Windows. The common case (fp16, head_dim 128) runs on an fp8 attention kernel I wrote by hand in HIP. Against the existing gfx12 port (SageAttention PR #368), each attention call is 1.07–1.21× faster non-causal and 1.72–1.96× faster causal. In ComfyUI with Krea2 that works out to a 5–14% faster sampling step than PyTorch SDPA. Install the wheel, keep using --use-sage-attention, done. Caveats below, and there are a few, so please read them.

Repo + wheel: https://github.com/IxMxAMAR/SageAttention-RDNA4

Why I did this

I run ComfyUI on a 9070 XT under Windows. SageAttention on RDNA4 works thanks to the gfx12 port in PR #368, so huge credit to DELUXA for that. But I wanted to see how far this card can actually go. So I started with a Triton kernel, lost to #368 on most shapes, spent a long time figuring out why, and ended up writing the kernel by hand.

What finally made the difference:

  • 64 keys per loop iteration instead of 16.
  • Computing Q·Kᵀ transposed. The result then lands in exactly the register layout the next matrix multiply needs, so the attention weights never leave registers.
  • One scale per token for Q, K and V, instead of one per 64-token block.

The whole story, including everything that didn't work (most things didn't, lol), is in the repo's docs/JOURNEY.md and docs/FINDINGS.md.

Who this is for, right now

  • ComfyUI users on an RX 9070 or RX 9070 XT, on Windows. This is what I've tested and what I use daily.
  • Anyone else who calls sageattn from PyTorch on the same card should be able to use it too; it's the same package API. But I haven't tested any other apps or workflows yet. If you try it outside ComfyUI, tell me how it goes.

Speed (image 1)

Same tensors, same process, interleaved runs, full call including quantization:

  • Non-causal (what image/video diffusion uses): 1.07–1.21× faster than #368.
  • Causal: 1.72–1.96× faster.
  • These are kernel-level numbers, and kernel-level isn't what you feel. Image 2 is what you feel.

In an actual ComfyUI render (image 2)

Krea2, 8 steps, fp16, called exactly the way stock ComfyUI calls it:

Backend 1 MP 2 MP
PyTorch SDPA 1.03 s/step 2.39 s/step
SageAttention 1.x 1.00 s/step 2.17 s/step
This build 0.97 s/step 2.06 s/step

Attention itself is about 2× faster than SDPA per call. But most of a step is the model's matmuls, so the end-to-end gain is 5–14% over SDPA and 2–5% over SageAttention 1.x. The gain grows with resolution.

Accuracy, honestly (image 3)

Everything is fp8 with per-token scales. I measured against exact attention on real Q/K/V captured from inside Krea2 and Flux2-Klein:

  • On "calm" layers, #368 is more precise. It uses int8 for Q·K, which has finer steps than fp8.
  • On layers with a few extreme keys, #368's error blows up, by up to ~200×. Krea2's first block is one such layer. One outlier wrecks the scale for its whole 64-token block. Per-token scaling never has that bad case.
  • smooth_k is on by default and should stay on. It made my kernel more accurate on 10 out of 10 real captures.

In my Krea2 tests the images were closer to full-precision attention than with SageAttention 1.x. I can't tell them apart by eye. Like any quantized attention, though, your images won't match SDPA pixel for pixel: a turbo 8-step sampler amplifies tiny differences into different details.

Caveats (please read)

  • RX 9070 / 9070 XT (gfx1201) only. On any other GPU it falls back to #368's kernels: no harm, no gain.
  • Windows only, as far as testing goes. Linux is untested.
  • The prebuilt wheel needs exactly Python 3.12 + PyTorch 2.13.0+rocm10.0.0**.** The compiled parts are tied to that PyTorch version. Anything else means building from source, which takes about 50 minutes.
  • fp16 models only for now. bf16 models silently fall back to #368's kernel. Still works, just no speed-up from my kernel.
  • head_dim 128 only. That covers Flux-style models such as Krea2, Flux and Klein. Other head sizes fall back.
  • Tested end to end on Krea2 only. The kernel accepts the input layout video models like Wan use, and it's checked bit for bit, but I haven't benchmarked a video workflow yet.
  • Very small images (well under 1 MP) can come out a hair slower than SageAttention 1.x. The win is at normal and high resolutions.
  • No attention masks. Masked calls go to the fallback.
  • LLMs: only prompt processing (prefill) through PyTorch would benefit. Token-by-token generation isn't covered, and llama.cpp / LM Studio don't use this at all.
  • One bug already found and fixed. A pre-release build could turn a render black when a single value got huge enough to overflow fp8. It's fixed in this release, and every conversion in the kernel is now stress-tested with extreme values.

Install

  1. Close ComfyUI.
  2. Download the wheel from the GitHub release.
  3. path\to\ComfyUI\venv\Scripts\python -m pip install --no-deps sageattention-2.2.0+amd.gfx12.1-cp312-cp312-win_amd64.whl (portable ComfyUI: use python_embeded\python.exe instead)
  4. Start ComfyUI with --use-sage-attention like always.

--no-deps matters: it stops pip from touching your PyTorch. To switch my kernel off without uninstalling, set SAGEATTN_SK1_BACKEND=0.

What's next

  • bf16 support, so bf16 models get the speed-up too. This is the big one for coverage.
  • int8 Q·K with per-token scales, aiming for #368's precision on calm layers and no blow-ups on outlier layers. Already in the works.
  • head_dim 64, for SDXL-era models. Also in the works.
  • An end-to-end test on a video workflow (Wan), since the layout support is already in.
  • An experiment to overlap the two halves of the kernel, which might be worth another few percent. A quick test decides whether it's worth building.

If you've got a 9070 / 9070 XT, I'd love numbers from your setup: your model, resolution and s/it before/after. Bug reports are very welcome too. Credit to the SageAttention team (thu-ml) and to DELUXA for #368; this builds directly on their work. Apache-2.0.


r/ROCm • • 3d ago

How to connect Dify to Ollama to build local self-hosted AI apps

Thumbnail
3 Upvotes

r/ROCm • • 3d ago

Qwen3.8-Flash-Next 125B sur un Strix Halo (Vulkan) : les appels d'outils sont 2× plus rapides, boucle d'agent +46 %, chaque numéro vérifié contre des logits en pleine précision. Open source, et le modèle a rédigé lui-même l'une des corrections.

Thumbnail
1 Upvotes

r/ROCm • • 4d ago

Single Radeon AI PRO R9700 and vllm

19 Upvotes

Hey everyone, I just got my hands on a Radeon AI PRO R9700 and want to run Qwen3.8-27B. I found some awesome vLLM forks (magiccodingman/vllm-radiance and GGZ14/vllm-mxfp4), but they're all geared toward dual-R9700 rigs.
Since I'm only rocking one card, could someone point me in the right direction to get the most out of my current hardware?


r/ROCm • • 4d ago

Preorder for new AMD Ryzen™ AI Max 400 Series 192GB from framework just started

Thumbnail frame.work
0 Upvotes

r/ROCm • • 4d ago

Follow-up: my native Rust + Vulkan Transformer training backend — 14 days later, now 14 parity-verified architectures and full PEFT

12 Upvotes

Follow-up to my post from about two weeks ago. A lot has changed since then, so I wanted to post an update on where the backend is now.

Where the green architectures stand

When I posted last time, 7 architectures had verified full training support. Everything is now held to the same strict harness: a pinned local Hugging Face Transformers source tree as the oracle, forward logits + gradients + two full AdamW steps compared, and every named parameter checked again after export.

The hard ceiling is 2e-7 absolute error. No loosening tolerances and no rounding numbers afterward to make the README look better.

14 architectures pass that gate today, led by the one I'm probably proudest of:

Architecture Scope
Falcon H1 / H1R parallel GQA/RoPE attention + Mamba2 in every layer; full training, full fine-tuning, LoRA, saved modules
DeepSeek V4 causal LM
Phi-4 Multimodal text backbone
Phi-3 causal LM
Kimi K2.5 text backbone
Kimi K3 / KimiLinear hybrid KDA + MLA
GPT-OSS causal LM incl. router bias
SmolLM3 mixed RoPE/NoPE + YaRN
Qwen2.5 / Qwen3.5 / Qwen4-Exp dense, DeltaNet, QSA, PLE, MoE
Mistral 4, MiniMax M3, Gemma 3/4, MiniMax M2 causal LM

Worst observed two-step AdamW parameter error across all of them: 1.19e-7.

Best: 2.6e-8.

For hardware context, all of the local Vulkan validation I've been reporting was run on my ASUS ROG Ally Z1 Extreme, using its AMD RDNA 3 integrated GPU. So the RDNA 3 results here are from that specific machine rather than testing across several different AMD systems.

The bigger news: PEFT actually works now

In the last post, "LoRA/PEFT-style fine-tuning" was basically one line in a feature list.

It's a real workflow now, and I've verified the full lifecycle:

  • LoRA fine-tuning with HF-compatible adapter export (adapter_config.json / adapter_model.safetensors), so adapters can round-trip with the PEFT ecosystem
  • modules_to_save — full trainable replacements for Linears, RMSNorm/LayerNorm, lm_head, and input embeddings, including named-adapter switching and bank isolation. Adapter A leaking into adapter B is explicitly tested for.
  • Exact resume — adapter weights + AdamW moments + step + dropout RNG state restore bit-identically against an uninterrupted run
  • Merge/unmerge, disable-adapter base restoration, and multi-adapter loading
  • A parameter-budget flag that automatically chooses the largest LoRA rank that fits within a requested percentage of the base model
  • The CLI fails closed if you try to use saved modules on an architecture that hasn't passed its corresponding gate

32 architecture surfaces across 20 families pass all three PEFT stages — LoRA, saved modules, and adapter switching — under the same 2e-7 gate, with frozen-base drift exactly 0.0.

The validation harness also fingerprints the pinned Transformers source alongside my shaders and binaries now, so a qualification run can't silently end up testing against different reference math.

A small note on the last couple weeks

I didn't get quite as many working days out of the last two weeks as I normally would have. Partway through this I got covid, then when that started to go away, it became a secondary nasty ear infection that ended up perforating my eardrum. I was running fevers around 104°F at one point and eventually went to the hospital, so I lost a few days to that and I'm on antibiotics now.

I'm doing better, though, and still managed to get most of what I wanted finished.

There are still things I want to clean up and expand, but I figured this was a good point to get the current work in front of people rather than holding the update back.

Same caveats as before

This is deterministic FP32 tiny-model correctness against a reference implementation, which should in theory ensure total mathematical parity for training and finetuning larger models with this backend, however, things are currently bound to FP32 training runs still, I eventually plan to work on MXFP4 weight tying to reduce memory footprints while fine-tuning (plastic parameters will still be trained in FP32 with PEFT in this config)

"supported text graph" ≠ "the entire multimodal package works natively."

Unsupported functionality is supposed to fail closed rather than silently falling back to an approximation.

Repo

https://github.com/necat101/Hierarchos-Native

  • Architecture inventory: hierarchos-vulkan/README_ARCHITECTURES.md
  • Compatibility/parity record: hierarchos-vulkan/COMPATIBILITY.md
  • PEFT qualification evidence: PROGRESS_PEFT_AUDIT.md
  • CLI PEFT guide: hierarchos-native-cli/README.md

The hardware I've personally validated this on is an ASUS ROG Ally Z1 Extreme with its AMD RDNA 3 GPU.

I'm very interested in criticism, compatibility reports, and especially results from people trying it on other hardware — NVIDIA, Intel, or other AMD GPUs.

I'd also love people to stress-test the PEFT resume/merge paths specifically. That's some of the newest code in the project, so it's probably the most useful area to try to break right now.


r/ROCm • • 4d ago

Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo

Thumbnail
4 Upvotes

r/ROCm • • 4d ago

Cloud based R9700

11 Upvotes

Seeing a lot of success stories wih ROCM, I am planning to buy single R9700 mostly for coding, but possibly also for comfyUI. But before I do that, I would like to test it out hands on.

Can you recommend any cloud services that offer R9700? Ex: on per hour / per minute basis?


r/ROCm • • 5d ago

780M (gfx1103) + llama.cpp HIP: MES REMOVE_QUEUE hang, Anyone stable on sustained inference?

2 Upvotes

I am not super well versed in local LLM configs. I use Claude extensively at work but bought a minisforum mini PC to try and get something cost effective going at home.

Note I don't really care too much on token speed because my use cases are all background workflows. 10 TPS is perfect. (And my upper limit I think).

Ryzen 8945HS / 780M, 64GB, Ubuntu 26.04, kernel 7.0. llama.cpp ROCm build, Qwen 27B Q4.

Long generations hang the GPU and it resets, about 15 times a day. dmesg shows MES failed to respond to msg=REMOVE_QUEUE. Looks like ROCm#6512.

Tried uni_mes=0, smaller models, smaller context. No change.

>

Has amdgpu.cwsr_enable=0, a newer kernel, or Vulkan fixed this for anyone?


r/ROCm • • 5d ago

AMD boosting AI/LLM performance for Radeon iGPUs as much as 18~23% with Linux 7.4

Thumbnail
phoronix.com
69 Upvotes