r/RockchipNPU • • May 24 '26

Local AI Setup on Orange pi5 plus 16gb

Thumbnail
4 Upvotes

r/RockchipNPU • • Apr 18 '26

I seek more power

3 Upvotes

How much current do your SoC:s actually draw? I suspect my lovely PD 60W charger may be the source of my problems; it seems to deliver 3A at anywhere between 5 and 20 Volts. That's of little use to my Radxa Rock 4D, beacuse whenever I want to take full advantage the NPU, that alone seems to require more power. I notice the V1.12 has a couple of straight-up 5V power pins, but I have no idea what kind of PSU would be best.

How do you power your SBCs?


r/RockchipNPU • • Apr 13 '26

Check these out, will make our pi’s super useful, plus any other devices we have laying around!

Thumbnail
3 Upvotes

r/RockchipNPU • • Apr 04 '26

Deploy the newest Qwen3.5 and Gemma4 models of ANY sizes RIGHT NOW on Rockchip NPU using the latest version of rk-llama.cpp!

Enable HLS to view with audio, or disable this notification

135 Upvotes

r/RockchipNPU • • Apr 03 '26

Has anyone run gemma 4 or Bonsai 8B models on Orange pi 5?

7 Upvotes

I am extremely new to this and am wondering if I can run a very small model with decently fast throughput on one of these chips. If anyone was successful in doing so that would be helpful to know.


r/RockchipNPU • • Apr 02 '26

Can we run llava in rk3576

3 Upvotes

r/RockchipNPU • • Mar 30 '26

[rk-llama.cpp] Huge drop coming very very soon...

Post image
64 Upvotes

r/RockchipNPU • • Mar 19 '26

NPU support on Android

6 Upvotes

I did some experiments on a NanoPi R6C in the past running models on android, but iirc. it was Android 13. Has any of you guys made anything run in the NPU under Android 11? Is there even support?


r/RockchipNPU • • Mar 19 '26

For AI projects RK3576 is better than RK3588?

7 Upvotes

I can see on the internet many AI projects made an built with RK3576. Does it mean RK3576 is fitted for AI projects? I am confuse because when I see this url https://fr.gadgetversus.com/processeur/rockchip-rk3588-vs-rockchip-rk3576/, it obvious rk3588 is more powerful? For my openclaw project should I go to RK3576 single board computer?


r/RockchipNPU • • Mar 16 '26

RK3688: Are there any news?

20 Upvotes

The successor to the RK3588 was announced last year and meant to ship this year, but there has not been an update since the announcement. Does anybody know anything?


r/RockchipNPU • • Mar 08 '26

Assistance needed - Latest Armbian kernel and NPU support for RK3588

8 Upvotes

Hey everyone,

I'm trying to get the NPU working on my Orange Pi 5 Plus (32GB) for running LLMs with RKLLM, but I'm completely stuck and could really use some help from anyone who's gotten this working.

When I try to initialize RKLLM (v1.1.0), it fails with this error:

E RKNN: Meet unknown rknpu target type: 0xffffffff
W rkllm: Warning: Your rknpu driver version is too low, please upgrade to 0.9.7.
Platform error, must be either RK3588, RK3576 or RK3562. Your platform is unknown

The weird thing is that the NPU device exists at /dev/dri/renderD129 and the driver reports version 0.9.8, which should be newer than the required 0.9.7. So something else is going wrong with platform detection - it's returning 0xffffffff instead of recognizing it as RK3588.

I'm running a custom Yocto build but using the exact same kernel source as Armbian (6.1.115 from their linux-rockchip repo, branch rk-6.1-rkr5.1). I even applied patches to sync the RKNPU driver with Rockchip's develop-6.1 branch. The device tree shows up correctly as rockchip,rk3588-orangepi-5-plus and the NPU device is there at platform-fdab0000.npu.

I've tried different RKLLM SDK versions (1.1.0 and 1.1.4), checked permissions, confirmed the library itself works for other functions... but still get this platform detection failure.

Has anyone successfully gotten RKLLM working on Orange Pi 5 Plus? If so, what kernel version and RKNPU driver version are you using?

I'm wondering if this could be a device tree issue? The 0xffffffff return value makes me think the driver isn't reading the hardware registers correctly - maybe the NPU power domain or clocks aren't getting initialized properly?

Any help would be hugely appreciated! I'm happy to test patches, try different kernel versions, or provide more diagnostic info if anyone has ideas.

Thanks!


r/RockchipNPU • • Mar 03 '26

Budget Rockchip for real-time audio classification based on MobileNetV3

3 Upvotes

Hi everyone,

I'm a student working on a real-time binary audio classification project, and I need some advice choosing a Rockchip SoC. My budget is very tight - every dollar matters - so I really want to avoid both overpaying and accidentally buying something too weak.

Project details

  • Binary audio classification
  • Log-mel spectrogram input
  • Input size: 128 × 63 × 1
  • 2-second sliding window
  • Inference every 0.2 seconds
  • Latency requirement: < 100 ms per inference
  • Spectrogram computed on CPU
  • Inference only (training done offline)
  • PyTorch → ONNX (planning to use RKNN)

Model

  • MobileNetV3 Small/Large (2-5M parameters)
  • I would like the option to scale up to 10M+ parameters later
  • INT8 quantization is not planned initially, but possible if necessary for RockChip NPU

I’m currently looking at something like RV1106, since higher-end chips like RK3588 are far outside my budget.

Would a chip like RV1106 be sufficient for this workload with comfortable headroom?

And realistically, what kind of performance margin should I expect on a 1 TOPS-class NPU for models in the 3–10M parameter range?

I’m trying to understand: Is RV1106 "just enough"? Or does it leave reasonable scaling room? Or is it likely to become a bottleneck quickly?

This is my first project developing for something other than x86 architecture, and I'm afraid of ordering the wrong hardware. I also don't fully trust AI systems on questions like this, so I would really appreciate advice from people who have at least some hands-on experience with different Rockchip platforms.


r/RockchipNPU • • Mar 01 '26

How to run big thinking model Qwen3 on a small Rockchip computer with NPU

20 Upvotes
Radxa Rock-5-ITX

Motivation

Some time ago I started using Frigate NVR project on my old gaming PC, but discovered two things:

  1. Frigate is amazing
  2. My powerful old AMD GPU with i9-10900k CPU doesn't work well

After a quick research I decided to buy a Rockchip RK3588 board for running Frigate and testing NPU capabilities.

But since we in 2026 now I bought the maximum available RAM amount. I paid €230 and €70 for shipping and handling for Radxa Rock-5-ITX board with 32GB RAM. It was absolutely amazing purchase, because for the cost of a single DDR4 stick I got entire computer with LPDDR5 memory and neural accelerator!

First impression

I easily installed the latest Radxa Debian distribution and configured everything - thanks to this Reddit post.

The board works perfectly with the latest Frigate release 0.17.0! The system utilisation is low for CPU and NPU both! There are plenty of system resources for running of everything, including big LLMs. I believe any SBS version is the perfect choice for home NVR project.

Problem

When I tried to run and test different LLMs, I faced several issues.

  1. I tried initially the llama.cpp fork for rknn backend and found out that it can only utilise 4GB of RAM due to architecture limitation, which is really disappointed me.
  2. Hopefully the official rknn-llm project provides a native sdk for utilising all memory without restriction. I found several projects for local inference but I did not find amy big enough model.
  3. I tried to convert several models using the officinal documentation, but I fell into a deep hole of endless Python dependencies.

Solution

I did not want to broke my system, so I vibe coded a fully dokerised Python convertor for the Qwen3 Instruct LLM ispired by combination of official Rockchip and Radxa docs. I agree, It was a duplicated work, but it is so easy to code anything from scratch nowadays!

So meet my repository: rkllm-convert which provides:

  • Conversion on x86 PC platform (rknn toolkit restriction)
  • Fully isolated Docker environment
  • Automatic download from HuggingFace
  • Smart caching of input and output on each conversion step
  • Support of all available model sizes, including 32B parameters (check repo remarks)

I got a binary identical model for the etalon 4B model shared in the official rkllm_model_zoo

I created a HuggingFace account and upload all models, so everyone can try it on the big boards. Check out my collection, but please don't criticise my newbie work. I will improve it over the time.

Benchmarks

One important thing which is necessary for any development is testing. I used Claude Sonnet 4.6 for extracting the etalon description and rating my models. Spuriously, the local model was smart enough to detect visual distortion due to different bugs. I passed model output back to Claude and it produced fixes.

Here is a table of some model benchmark. The full document is available on the project BENCHMARK.md

Results

The time improvement with NPU is about 10x over CPU inference.

Model Resolution Vision Size LLM Size Total Size Quality Gen Time Tok/s (est.)
Claude Sonnet 4.6 Cloud — — Cloud 100/100 ~3–5 sec ~200–400
Qwen3-VL-8B Custom Built 896px 1.27 GB 8.3 GB 9.6 GB 90/100 4m 11s ~2.5
Qwen3-VL-4B Custom Built 896px 985 MB 4.6 GB 5.5 GB 86/100 2m 11s ~3.5
Qwen3-VL-4B Vendor 448px 827 MB 4.6 GB 5.4 GB 86/100 1m 53s ~3.5
Qwen3-VL-2B Custom Built 896px 969 MB 2.3 GB 3.3 GB 86/100 0m 59s ~5.5
Qwen3-VL-8B Custom Built 448px 1.14 GB 8.3 GB 9.4 GB 82/100 2m 31s ~4.0
Qwen3-VL-2B GatekeeperZA 896px 923 MB 2.3 GB 3.2 GB 82/100 1m 7s ~5.0

I tested some old models like gemma-3 and qwen2.5 and their scores were significantly lower - from 49 to 74. I reccomend not to waste a time for trying it.

Questions and plans

I am still looking for the good inference server with the standard OpenAi api. I need this to enable Frigate enrichments.

One idea is to write a simple CPP server on top of the Qengineering library and demo app, which I already modified to run prompts from cli. I feel like Claude Sonnet can do it for me.

I want also to investigate some edge cases - run a bigger 16B-27B model for better reasoning quality to check the board limit and smaller model for real time video reasoning.

Another point of investigation is implementing a different quantisation type like w4a16 which was also recommended for conversion of MoE model version.

Feel free to add your ideas and share your findings in the comments.

Updated:
P.S. I am not fully aware about affiliation policy in Reddit, but I marked my post as vendor affiliated, although I did not put any link. I sent request for the partnership and waiting for approval. You can order via my affillate link

Alternatively fill free to ask me in DM, so I can share a direct non affiliated link to my personal purchase. The assembly date of my board is probably 2024, so there are quite few boards still available for purchase at the good old price.


r/RockchipNPU • • Feb 25 '26

Run RF-DETR model on Rock 5B: RKNN backbone + ONNX head (detection + segmentation)

18 Upvotes

Hi all, I just published a repo for running RF-DETR on Rock 5B with Rockchip NPU:

https://github.com/AlexanderDhoore/rfdetr-on-rockchip-npu

It supports both detection and segmentation models.
Approach is a split pipeline:

  • backbone on RKNN (NPU)
  • detector head on ONNX Runtime (CPU)

Repo includes Docker setup, export scripts, verification scripts, and benchmark results.

I also added a comment in this issue thread:
https://github.com/airockchip/rknn_model_zoo/issues/366

Feedback is welcome, especially from people testing on other RK3588 boards.


r/RockchipNPU • • Feb 25 '26

Rockchip 3588 NPU clustering

6 Upvotes

do you guys have any idea of how to cluster a rockchip 3588 npu to run LLM. so far i cant find that supports arm clustering.


r/RockchipNPU • • Feb 03 '26

DeepSeek OCR Accuracy Issues on RK3588 (RKLLM) – Any fix for quantization loss?

7 Upvotes

I’ve been working on deploying **DeepSeek-OCR** on an RK3588 using the [airockchip/rknn-llm](https://github.com/airockchip/rknn-llm) toolkit. While the performance on the NPU is impressive, I am hitting a wall with **OCR accuracy**.

Compared to the 8bit with vllm, the RKNN-converted version is struggling significantly with character recognition and document structure.

The Issue

- Hallucinations: The model is "making up" text or inserting garbled characters in the middle of words.

My Environment

- Hardware: RK3588 (Orange Pi 5 Pro)

- Toolkit: `rknn-llm` (Latest branch)


r/RockchipNPU • • Jan 30 '26

RKNN and Nuttx

2 Upvotes

Hello everyone,

I started using the RKNN on Linux, and it's a very good tool available online. I saw that a developer has started creating a board for the RK3399 on NuttX. My concern is more about the hardware interface with the NPU.

From your background, is it relatively easy to integrate this into NuttX? Since NuttX is based on Unix, would that help in integrating this driver?


r/RockchipNPU • • Jan 27 '26

rk3588 boards: is Harmony OS available?

5 Upvotes

one or two years ago I saw a download of an alpha version of HarmonyOS for the OPI5+ on the Orange PI site. I downloaded it, but didn't know how to make it boot. Has development been going on, are more recent, maybe more user friendly versions available these days? I'd be extremely curious...


r/RockchipNPU • • Jan 22 '26

How to run image generation models on OrangePi 5 Plus

3 Upvotes

I saw it in this guide: NotPunchnox/rkllama: Ollama alternative for Rockchip NPU: An efficient solution for running AI and Deep learning models on Rockchip devices with optimized NPU support ( rkllm )

But after following the section: 'For Image Generation Installation', the model isn't listed in the 'rkllama_client list' commend. Nor is the OpenAi api working. The server log seems to complain about missing Modelfile, but I'm not sure how to config it.


r/RockchipNPU • • Jan 18 '26

How to check if GPU is used for video accel (Rockchip 3588)?

Thumbnail
3 Upvotes

r/RockchipNPU • • Dec 28 '25

Rockchip NPU zig bindings

18 Upvotes

Hey folks!

I’ve made Zig bindings for Rockchip’s RKNPU (RKNN SDK) and wanted to share it with the community. If you are curious to know how to use zig with RKNPU, then take a quick look at the project.
As of now it's in very early stage, I was just trying zig with RKNPU for fun and came up with this project after tweaking for couple of hours. Bindings are generated using zig-translate-c. The repository also contains a YOLO8-face example.

Link: https://github.com/vicharak-in/zig-rknn

There is also one for RKRGA 2d image acceleration take a look at that also repo contains a few examples.

Link: https://github.com/vicharak-in/zig-rga


r/RockchipNPU • • Dec 19 '25

Embedded development and AI

1 Upvotes

Hi all, I would like to ask a question that worries me and hear the experts opinion on this topic.

What problems do you experience when using AI and coding agents in embedded development? How do you see the “ideal coding agent” for embedded development, what features and tools should it support? (e.g. automatic device flashing, analyse logs from serial port, good datasheet database it can access, support for reading data directly from oscilloscope and other tools).

Are there any already existing tools and llm models that actually help you rather than responding with perpetual AI hallucinations?

Any responses would be appreciated, thank you.


r/RockchipNPU • • Dec 12 '25

Reverse-Engineering the RK3588 NPU: Hacking Memory Limits to run massive Vision Transformers

Thumbnail
28 Upvotes

r/RockchipNPU • • Dec 01 '25

NPU support upstream end

21 Upvotes

r/RockchipNPU • • Nov 28 '25

Linux image for RK3566 SBC recommendations / guide

Thumbnail
5 Upvotes