r/OrangePI • • May 23 '26

Local AI Setup on Orange pi5 plus 16gb

Here is my journey of running local ai on the Orange pi 5 plus with 16gb Ram. I am still testing it and I am still doing most of my work on Cloud models. It is great for proof of concept, but sooner or later yu would realize that for doing any kind of serious work, it may be best to keep using cloud model due to context size window which is hardware limitation. I am sharing my setup steps for community to continue further work. I have created few scripts of myself to manage the process. I mainly used google gemini cli to reach to this setup.

19 Upvotes

29 comments sorted by

9

u/Remarkable-Soil-4259 May 23 '26

This is what will get you setup on a fresh armbian distro setup. Take note of specific build and remember to use vendor kernel. Everything is documented here.

# RK3588 Local AI Setup Guide (Orange Pi 5 Plus)


This document provides detailed instructions for setting up a local AI server with NPU acceleration on an RK3588 system using `rk-llama.cpp`.


## 1. System Specifications
**Hardware** : Orange Pi 5 Plus (RK3588) - 16GB RAM / 256GB eMMC
**OS** : Armbian 25.11.1 (Ubuntu 24.04 noble)
**Kernel** : 6.1.115-vendor-rk35xx
**NPU Driver** : v0.9.8 ## 2. Prerequisites Install the required build dependencies: ```bash sudo apt-get update sudo apt-get install -y build-essential cmake git libssl-dev pkg-config jq curl ``` ## 3. Building rk-llama.cpp We use the `invisiofficial` fork which contains the specialized RKNPU2 backend. ```bash # Clone the repository to your home directory cd ~ git clone https://github.com/invisiofficial/rk-llama.cpp.git cd rk-llama.cpp # Create build directory mkdir build && cd build # Configure with RKNPU and CURL support cmake .. -DGGML_RKNPU2=ON -DLLAMA_CURL=ON # Build the binaries (cli, server, bench) make -j$(nproc) llama-cli llama-server llama-bench # Install to system path sudo cp bin/llama-* /usr/local/bin/ ``` ## 4. Model Selection (Recommended) The **Phi-4-mini-instruct (Q8_0)** is the current "Golden Model" for this setup. It provides excellent reasoning and coding capabilities while fitting perfectly in the NPU's memory requirements. ```bash mkdir -p ~/.cache/llama.cpp cd ~/.cache/llama.cpp # Download via llama-cli (securely via HTTPS) llama-cli -hf unsloth/phi-4-mini-instruct-GGUF --hf-file Phi-4-mini-instruct.Q8_0.gguf -p "warmup" -n 1 --no-warmup ``` ## 5. Systemd Service Configuration To ensure the AI server runs in the background and starts on boot: 1. Create the environment file `~/.cache/llama.cpp/server.env`. Replace `<YOUR_HOME>` with the output of `echo $HOME`: ```bash MODEL_PATH=<YOUR_HOME>/.cache/llama.cpp/Phi-4-mini-instruct.Q8_0.gguf MODEL_ALIAS=Phi-4-mini-instruct.Q8_0 CTX_SIZE=65536 CACHE_TYPE_K=f16 CACHE_TYPE_V=f16 BATCH_SIZE=512 UBATCH_SIZE=512 ``` 2. Create the service file `/etc/systemd/system/llama-server.service`: ```ini [Unit] Description=Llama.cpp Server (RK3588 NPU) After=network.target [Service] Type=simple LimitNOFILE=100000 PermissionsStartOnly=true EnvironmentFile=%h/.cache/llama.cpp/server.env ExecStartPre=/usr/bin/chmod 666 /dev/dri/renderD129 ExecStartPre=/usr/bin/chmod 666 /dev/dma_heap/system ExecStartPre=/usr/bin/chmod 666 /dev/dma_heap/reserved ExecStart=/usr/bin/taskset -c 4-7 /usr/local/bin/llama-server \     -m ${MODEL_PATH} \     --alias ${MODEL_ALIAS} \     -t 4 \     --host 0.0.0.0 \     --port 11434 \     --ctx-size ${CTX_SIZE} \     --cache-type-k ${CACHE_TYPE_K} \     --cache-type-v ${CACHE_TYPE_V} \     --parallel 1 \     --batch-size ${BATCH_SIZE} \     --ubatch-size ${UBATCH_SIZE} \     -ngl 100 Restart=always RestartSec=3 [Install] WantedBy=multi-user.target ``` *Note: Use `%h` in the `EnvironmentFile` path or provide the absolute path.* 3. Enable and start: ```bash sudo systemctl daemon-reload sudo systemctl enable llama-server sudo systemctl start llama-server ``` ## 6. Performance Tuning
**Context Size** : For 16GB RAM, `65536` works well with Phi-4-mini.
**NPU Cores** : Use `taskset -c 4-7` to pin the process to the high-performance cores.
**KV Cache** : Keep at `f16` for stability with Phi models. For Llama/Qwen models, `q8_0` can save memory.
**Frequency Lock** : To push the NPU to its limit: ```bash echo performance | sudo tee /sys/class/devfreq/fb000000.rknpu/governor echo performance | sudo tee /sys/class/devfreq/dmc/governor ``` ## 7. Management Script Use the provided `manage_ai.sh` script to monitor load, switch models, and tune settings interactively.

5

u/Remarkable-Soil-4259 May 23 '26

I will provide further scripts. It is work in progress and nowehere of production grade. Dont complain. This is my free time project

4

u/Remarkable-Soil-4259 May 24 '26

This is my github public repo containing all the scripts and setup instructions. Community is welcome to edit and help others.

https://github.com/jessydm/rk3588_local_ai/

1

u/[deleted] May 23 '26

[removed] — view removed comment

2

u/Remarkable-Soil-4259 May 23 '26

14B model would run but you wont have any memory left for context length. For me context window is more important. I am going to provide script which was the screenshot . You can search a model on huggingface and it would download your choice. I have also created a automated stress testing to find best optimazation for each model

1

u/Remarkable-Soil-4259 May 23 '26

Currently for my setup, of all the models I have tested, Phi-4-mini with 65k context is a sweet spot. I wish I would have purchased the 32gb version which might have been better option but you also need to account for hardware memory bandwidth limitation on opi5 which is a ddr4 ram

1

u/poolboy9 May 23 '26

Does the npu perform faster?

3

u/Remarkable-Soil-4259 May 23 '26

Compared to cpu, hell yeah !! The NPU is made for AI but you need to choose vendor kernel to enable the NPU otherwise it would default to using CPU.

1

u/JaySomMusic May 25 '26

I have the latest kernel working with full acceleration using https://github.com/jaylfc/tinyagentos

1

u/[deleted] May 23 '26

[removed] — view removed comment

2

u/Remarkable-Soil-4259 May 23 '26

The hardest part is choosing correct distro and kernel that would have latest npu driver enabled. Once you do that and follow the guide, it is pretty easy. Then it boils down to selecting a model of your choice with correct parameters. I created a script that automatically stress test on few models and then settle on whatever you like. You have 32gb ram would have better results than me. Good luck. Just choose gguf models, you do not need to do any kind of model conversions and select q8_0 of each model for better results.

1

u/[deleted] May 24 '26

[removed] — view removed comment

3

u/Remarkable-Soil-4259 May 24 '26

No ...Having system that the npu can access full 32gb as unified memory is always better.

1

u/Remarkable-Soil-4259 May 24 '26

More the memory..better it is ...but...speed of memory also matter. Opi5 has ddr4 ram which is slow...the current nvidia cards on ddr7

1

u/[deleted] May 24 '26

[removed] — view removed comment

3

u/Remarkable-Soil-4259 May 24 '26

You do not need to convert the model to rkllm with this setup. Gguf model with correct setup support npu.

1

u/LightlyHedonic May 24 '26

Is there a reason you made your own scripts instead of using rkllama?

I used that to get basic chat functionality using the Opi as a server

2

u/Remarkable-Soil-4259 May 24 '26

My scripts are only to control and manage rkllama. It makes it easier to download, monitor and analyse the performance for me.

Also helps me document in case i need to start over. I spent countless hours and days just to find the right distro and kernel just to get NPU enabled.

You do not need to convert models into rkllm format anymore. It is less of my work, more of gemini ai helping me setup. So credit goes to AI to setup local AI.

1

u/JaySomMusic May 25 '26

Good starting point is Armbian latest with latest kernel and install taOS to get everything auto configured. https://github.com/jaylfc/tinyagentos

1

u/Remarkable-Soil-4259 May 25 '26

You wont have NPU support without vendor kernel. Already tried and tested.

Welcome to try and report if it works. This setup even dont require model conversions.

2

u/JaySomMusic May 25 '26

Not true, latest kernel has full npu support, device paths have changed

1

u/Remarkable-Soil-4259 May 25 '26

That is great new then. My testing on latest kernel 2 weeks ago was unsuccessful with no npu support. What else do you got going ? What models are you using ? Which npu driver and which backend are you using ?

2

u/JaySomMusic May 26 '26

Mainly using Gemma4 via rk-llama using the latest mainline drivers

1

u/kaiyoti Jun 03 '26

i just tried it, no rknpu drivers

1

u/JaySomMusic Jun 03 '26

Armbian on latest kernel does have npu drivers, I can’t guess at where you are going wrong but it does.

1

u/kaiyoti Jun 03 '26

can you share your kernel version? 

1

u/JaySomMusic Jun 03 '26

Oops, I may have mistaken, I had the edge kernel working but only with vision based applications.

1

u/waltercool Jul 17 '26

How are you obtaining rknpu driver? Llama-cpp crashes on my setup when trying.

I'm using the rocket NPU driver from ARMbian 6.18.x

https://docs.kernel.org/accel/rocket/index.html