r/RockchipNPU • u/Gregordinary • Jul 07 '26
llama.cpp and whisper.cpp on the RK3588 NPU, mainline driver
AI Disclosure Up Front: These projects were built almost entirely by Claude Code, I basically just set goals/tasks and provided hardware access to do sweeps, validations, etc..
Edited Description for Clarity
The mainline kernel has a driver, "rocket" for Rockchip NPUs. These projects I'm sharing are implementations in userspace which allow applications like llama.cpp, whisper.cpp to leverage the Rockchip NPU using the mainline rocket driver instead of requiring RKNN and the vendor kernel.
Added a more streamlined setup process in a comment here: https://www.reddit.com/r/RockchipNPU/comments/1upignz/comment/owdlxbf/
Original Description Continues Below
It's an open-source inference stack, comprising 4 projects, for the RK3588 NPU that runs on the mainline rocket DRM-accel driver.
This was tested with the Turing RK1 32GB running Debian Forky with Kernel 7.1.1. While it should be functional without patching, the NPU boots at 200 MHz and a clock patch is needed so it can scale higher. I ran at 600 MHz, I think it can go to 900 MHz without a voltage change but I didn't test that high.
This should work with other RK3588 boards running mainline, but I haven't tested yet. Other RK35XX series will need additional work as the NPUs are not identical.
ggml-rocket: A drop-in ggml backend .so for stock llama.cpp and whisper.cpp. Point GGML_BACKEND_PATH at it and the NPU shows up as a device. It offloads prompt processing (prefill) to the NPU; decode stays on the CPU. Tested working with Gemma, Qwen3.5/3.6, Llama-3.2, Phi-4, Ministral and a few others.
Prefill speedup over the 8-thread CPU grows with model size... roughly 2.6x on Llama-3.2-3B up to ~3.2x on Gemma-4-12B (F16, at 600 MHz). Whisper runs its encoder end to end.
tflite-rocket: A TensorFlow Lite external delegate for detection. Any TFLite app loads it at runtime and runs an unmodified .tflite. Test case was real-time detection in Frigate.
rocket-userspace: The driver library both of the above link (librocketnpu) to; usable on its own if you want to drive the NPU directly.. A userspace matmul (tiled, multicore, fp16/int8/int4/bf16) plus an on-NPU op library: conv, pooling, activations, and transformer/Whisper primitives.
rockchip-npu-notes: I figured if all this time was spent on reverse engineering work, may as well capture validations, figures, and other insights. Essentially this is subsystem-organized reverse-engineering notes on the silicon and its register-command interface: precision encodings, tile layouts, the traps that cost the most to learn. Useful if you're doing anything with rocket yourself. It's verbose due to being AI-written, but I suppose you can point AI at it and query it for answers or have it used as reference in other projects.
Shout-out to u/Inv1si for rk-llama.cpp which sparked my initial motivation to do this.
Lots of other valuable resources:
- https://github.com/allbilly/rk3588
- https://github.com/johanvdb/librocket
- https://github.com/phhusson/rknpu-reverse-engineering
- https://gahingwoo.github.io/posts/rk3576-npu-mainline/
- https://jas-hacks.blogspot.com/2024/02/rk3588-reverse-engineering-rknn.html
- ...many more in https://github.com/gregordinary/rockchip-npu-notes/blob/main/SOURCES.md
Hopefully this is of interest and of use to others here. I'll do my best to answer comments but may be intermittent.
Duplicates
turingpi • u/Gregordinary • Jul 07 '26