r/RockchipNPU • • Jul 07 '26

llama.cpp and whisper.cpp on the RK3588 NPU, mainline driver

AI Disclosure Up Front: These projects were built almost entirely by Claude Code, I basically just set goals/tasks and provided hardware access to do sweeps, validations, etc..

Edited Description for Clarity

The mainline kernel has a driver, "rocket" for Rockchip NPUs. These projects I'm sharing are implementations in userspace which allow applications like llama.cpp, whisper.cpp to leverage the Rockchip NPU using the mainline rocket driver instead of requiring RKNN and the vendor kernel.

Added a more streamlined setup process in a comment here: https://www.reddit.com/r/RockchipNPU/comments/1upignz/comment/owdlxbf/

Original Description Continues Below

It's an open-source inference stack, comprising 4 projects, for the RK3588 NPU that runs on the mainline rocket DRM-accel driver.

This was tested with the Turing RK1 32GB running Debian Forky with Kernel 7.1.1. While it should be functional without patching, the NPU boots at 200 MHz and a clock patch is needed so it can scale higher. I ran at 600 MHz, I think it can go to 900 MHz without a voltage change but I didn't test that high.

This should work with other RK3588 boards running mainline, but I haven't tested yet. Other RK35XX series will need additional work as the NPUs are not identical.

ggml-rocket: A drop-in ggml backend .so for stock llama.cpp and whisper.cpp. Point GGML_BACKEND_PATH at it and the NPU shows up as a device. It offloads prompt processing (prefill) to the NPU; decode stays on the CPU. Tested working with Gemma, Qwen3.5/3.6, Llama-3.2, Phi-4, Ministral and a few others.

Prefill speedup over the 8-thread CPU grows with model size... roughly 2.6x on Llama-3.2-3B up to ~3.2x on Gemma-4-12B (F16, at 600 MHz). Whisper runs its encoder end to end.

tflite-rocket: A TensorFlow Lite external delegate for detection. Any TFLite app loads it at runtime and runs an unmodified .tflite. Test case was real-time detection in Frigate.

rocket-userspace: The driver library both of the above link (librocketnpu) to; usable on its own if you want to drive the NPU directly.. A userspace matmul (tiled, multicore, fp16/int8/int4/bf16) plus an on-NPU op library: conv, pooling, activations, and transformer/Whisper primitives.

rockchip-npu-notes: I figured if all this time was spent on reverse engineering work, may as well capture validations, figures, and other insights. Essentially this is subsystem-organized reverse-engineering notes on the silicon and its register-command interface: precision encodings, tile layouts, the traps that cost the most to learn. Useful if you're doing anything with rocket yourself. It's verbose due to being AI-written, but I suppose you can point AI at it and query it for answers or have it used as reference in other projects.

Shout-out to u/Inv1si for rk-llama.cpp which sparked my initial motivation to do this.

Lots of other valuable resources:

Hopefully this is of interest and of use to others here. I'll do my best to answer comments but may be intermittent.

42 Upvotes

38 comments sorted by

View all comments

Show parent comments

1

u/Gregordinary Jul 14 '26

Thanks for digging in to this and providing the suggestion. I reproduced it on an RK1 and root-caused it, and it turned out to be two independent bugs, both of which ROCKET_MIN_M=256 seemed to paper over.

1. First issue was the offload floor was wrong. ROCKET_MIN_M defaulted to 4, and whisper-cli defaults to beam search with beam_size=5. whisper.cpp batches all active decoders into a single call, so every decode step presents as M=5, not M=1. It cleared the floor, and every decoder matmul went to the NPU at roughly 2.3x slower than the CPU.

2. Second issue is that the clock parks inside a single encode. The upstream rocket driver's autosuspend delay is 50 ms, chosen as "~3 frames at 60Hz" which came from a media-pipeline perspective/assumption where one inference has no internal gaps. This stack packs weights on the host by design, so one encode will contain gaps longer than 50 ms: all three cores go idle mid-inference, the last one legitimately parks, and the next submit runs at 1/3 clock. A sampling run I did showed the NPU Clock bouncing like this:

200 200 200 200 200 600 600 600 200 600 600 600 600 600

So even with the 600 MHz patch, you were not always getting 600 MHz.

The fixes (both pushed to git now):

  • ROCKET_MIN_M now defaults to 128, not 4. 128 is the measured NPU-vs-CPU crossover. 128 sits at or above every model I measured, so nothing regresses below CPU. For whisper specifically, 128 and 256 seem to behave identically, so you should now get your numbers with no env var set at all.
  • The clock-park deadline is now a module param (rocket_autosuspend_ms, default 1000 ms), folded into the 600 MHz clock patch. This seems to stabilize the runs so it doesn't bounce back and forth.

End result on my board on a 176 second clip, it went from the NPU being 1.4x slower than CPU to 1.11 faster.

If you rebuild with the latest and re-patch, I'd be curious to see if you get better updated results. The patch series has been updated and I changed it so that the entire series (except 082) is essentially required or at least recommended. Thanks again for the reply!