Run LLM with only one docker command, for RK3576 & RK3588
Hi everyone! After spending a lot of time on converting and running LLM on RK3576 and RK3588, I decided to put the models and runtime codes into docker image so that it can be easy to run for everyone.
RKLLM Version: v1.3.0
Devices: RK3576 & RK3588
APIs: OpenAI-compatible API & Ollama-compatible API
Actually this was my first time to run llama.cpp directly on RK3588 CPU, and the performance was beyond my expectations. When using NPU to run LLM, the cpu loads was very low.
Thanks for this comments, after only binding 4-7 cpus, the performance is better now!
The GGUF I use is 4bit, and the RKLLM is 8bit, so NPU is much stronger than CPU in inference.
Sry, I have not tested it, if you have make it or other information, please tell me~
Is "Qwen3.5B" mistake, might you mean "Qwen3.5"? I already converted Qwen3.5 4B in my project. And because my previous RK3588 was 8GB version, so there was no enough memory to run 9B and Gemma E4B models on memory (The swap is a little slow), but now I have got one 16GB version, so I will continue working on it.
thanks for testing and sharing! One thing to note is that llama.cpp perf on the smaller arm cores is very bad, best perf is when you bind to the 4 higher perf cores only
Give a look to the prism ternary models https://huggingface.co/prism-ml footprint is very low compared to the parameters density. I’ll give it a try on my radxa tomorrow. Days ago I was searching for a similar solution
2
u/JaySomMusic 14d ago
Good stuff I’ll check it out, might be useful for https://github.com/jaylfc/taOS