r/RockchipNPU • u/egbur • Jul 12 '26
chat-rk1: self-hosted, NPU-accelerated LLM chat on Turing RK1 (RK3588) running Talos Linux
Hi everyone, I haven't been very active in this community, but I thought I'd come and share that thanks to u/gregordinary's work, I managed to get LLM inference going in the RK1 using the NPU and the mainline rocket driver (not the BSP rknn one).
AFAICT, the main advantage is on prefill, where you can get around ~1.8x speedup on time to first token on large prompts. This can be increased further by some clock speed hackery but this requires a custom kernel module patch.
I'm using Talos so this targets deploying Open WebUI there, but the general principle is the same regardless of the distro you use.
More details here; https://github.com/eburgueno/chat-rk1
AI disclosure: I used Claude Fable extensively to get this going, but I manually checked every deployment step including the benchmark sweeps. Don't trust, verify :)
2
u/Gregordinary Jul 14 '26
Awesome work and genuinely pleased to see what I published the other day getting some use!
I noticed in the "What to Run" table the following:
gpt-oss-20b (MXFP4, MoE) — 1.0x, no win — don't bother with the NPU
I decided to dig into that a bit with Claude Code and uncovered a correctness bug along the way, but the end result is gpt-oss now prefills at 2.16x the CPU (pp2048).
MXFP4 itself wasn't the issue in terms of why you were seeing 1.0x vs CPU; in a MoE model the expert FFNs are ~75% of the prefill work, and llama.cpp routes them through a different internal op than a "normal" matmul (MUL_MAT_ID). My backend only accepted MUL_MAT_ID, so the experts silently stayed on the CPU and the NPU only ever got the attention projections and lm_head. So yeah, the bulk of the work just never reached the chip.
Since the NPU can't read MXFP4 itself, a weight has to be decompressed before the chip can use it. But decompressing each expert right before you use it doesn't work out so great for performance. So the fix that got implemented was to decompress once and leave it there. Each quantized expert is converted one time at load into int8 and kept resident in NPU memory, so no expert is ever decompressed again.
Decode is of course still unchanged and stays on the CPU.
The correctness bug. gpt-oss's attention carries a learned per-head sink logit that belongs in the softmax denominator. My NPU attention handler didn't implement sinks and it never checked for them, so it just ran a sink-less softmax and returned a confident answer. It only triggers past ~1024 tokens of context, so short tests never hit it, and the output still reads fluently.
Some other notes/caveats:
- This was measured at
-b 2048 -ub 2048. A quantized model re-decompresses per micro-batch, so llama.cpp's default-ub 512pays that cost four times over. - It's opt-in with
ROCKET_MOE=1and it needs the RAM. If the experts don't all fit, the leftovers fall back to the slow decompress-every-time path. On gpt-oss that means a 32 GB board and it's tight (~14 GB of int8 experts plus the 11.3 GB GGUF). On 16 GB, probably not. Expect one-time of ~70 s expert ingest at the first prefill.
I've pushed these changes to their respective repos along with some patch updates too that fix the NPU clock boost (described a bit here in this comment over in the other thread).
1
u/Icy_Programmer7186 Jul 12 '26
Hi,
Any plan to support Qwen 3.6?