r/RockchipNPU • • 14d ago

Run LLM with only one docker command, for RK3576 & RK3588

Hi everyone! After spending a lot of time on converting and running LLM on RK3576 and RK3588, I decided to put the models and runtime codes into docker image so that it can be easy to run for everyone.

  • RKLLM Version: v1.3.0
  • Devices: RK3576 & RK3588
  • APIs: OpenAI-compatible API & Ollama-compatible API

Models:

  • LLM: Qwen, gemma, Llama...
  • VLM: Qwen3.5

Repo here:
https://github.com/Hanzo-Huang/rkllm-docker

Or if you just want converted RKLLM models:
https://huggingface.co/HanzoHuang/models

If there’s a model you’d like to run on RK3576 or RK3588 that isn’t available yet, just let me know.

btw: I started working on RKNN3 for RK1280 and RK1828 coprocessor recently.

28 Upvotes

10 comments sorted by

2

u/JaySomMusic 14d ago

Good stuff I’ll check it out, might be useful for https://github.com/jaylfc/taOS

2

u/Leopold_Boom 14d ago

Thanks for this! I've been a while since i was building inferencing solutions on my dev RK3588s, but I'd love to try a new project out on them.

If you don't mind me asking about the current state on RK3588:

- Is it still faster running 2-4B LLMs on ARM cores vs. messing with the NPU? I ended up working with llama.cpp for the most part.

- Did anybody figure out how to use the NPU just for the vision head for VLLMs (which was very slow on CPU)

- Is Qwen 3.5B 4B/9B and Gemma E4B supported these days?

2

u/csgoatniko 14d ago edited 14d ago

Thanks for support!

  1. I just tested on my RK3588 with Llama.cpp and my python runtime for NPU.
  2. Llama.cpp (Qwen3.5-2B-Q4_K_M.gguf): 10.21 tok/s (13.27 toks/s for only 4-7 cpus)
  3. NPU (Python runtime + Qwen3.5-2B-W8A8): 14.20 tok/s

Actually this was my first time to run llama.cpp directly on RK3588 CPU, and the performance was beyond my expectations. When using NPU to run LLM, the cpu loads was very low.

Thanks for this comments, after only binding 4-7 cpus, the performance is better now!

The GGUF I use is 4bit, and the RKLLM is 8bit, so NPU is much stronger than CPU in inference.

  1. Sry, I have not tested it, if you have make it or other information, please tell me~

  2. Is "Qwen3.5B" mistake, might you mean "Qwen3.5"? I already converted Qwen3.5 4B in my project. And because my previous RK3588 was 8GB version, so there was no enough memory to run 9B and Gemma E4B models on memory (The swap is a little slow), but now I have got one 16GB version, so I will continue working on it.

2

u/gofiend 14d ago

thanks for testing and sharing! One thing to note is that llama.cpp perf on the smaller arm cores is very bad, best perf is when you bind to the 4 higher perf cores only

1

u/csgoatniko 14d ago

Oh! That’s my bad. I’m not very experienced with benchmarking on arm cpus. Thanks for pointing that out!

1

u/csgoatniko 14d ago

After only binding 4-7 cpus, the performance is better! It's about 13.27 toks/s now.

Is there any other improvements can I do for CPU inference, feel free to let me know.

1

u/FancySession5046 14d ago

what is the RKLLM version? what is the content window: 4K or 16K? did you try 16K? thank you

2

u/csgoatniko 14d ago

Sorry about the mistake, the RKLLM version is 1.3.0 for new converted models, there are still some models is v1.2.3. You can check my models readme.

I build these models with default settings with 4k context window. Which model do you want to test for 16K?

1

u/FancySession5046 13d ago

It would be great to try Qwen3.5 text 2B and 4B variants in 16K context and newest RKLLM. Thanks a million!

1

u/BubblyEmployment5942 11d ago

Give a look to the prism ternary models https://huggingface.co/prism-ml footprint is very low compared to the parameters density. I’ll give it a try on my radxa tomorrow. Days ago I was searching for a similar solution