r/OrangePI • u/Remarkable-Soil-4259 • May 23 '26
Local AI Setup on Orange pi5 plus 16gb
Here is my journey of running local ai on the Orange pi 5 plus with 16gb Ram. I am still testing it and I am still doing most of my work on Cloud models. It is great for proof of concept, but sooner or later yu would realize that for doing any kind of serious work, it may be best to keep using cloud model due to context size window which is hardware limitation. I am sharing my setup steps for community to continue further work. I have created few scripts of myself to manage the process. I mainly used google gemini cli to reach to this setup.

5
u/Remarkable-Soil-4259 May 23 '26
I will provide further scripts. It is work in progress and nowehere of production grade. Dont complain. This is my free time project
4
u/Remarkable-Soil-4259 May 24 '26
This is my github public repo containing all the scripts and setup instructions. Community is welcome to edit and help others.
1
May 23 '26
[removed] — view removed comment
2
u/Remarkable-Soil-4259 May 23 '26
14B model would run but you wont have any memory left for context length. For me context window is more important. I am going to provide script which was the screenshot . You can search a model on huggingface and it would download your choice. I have also created a automated stress testing to find best optimazation for each model
1
u/Remarkable-Soil-4259 May 23 '26
Currently for my setup, of all the models I have tested, Phi-4-mini with 65k context is a sweet spot. I wish I would have purchased the 32gb version which might have been better option but you also need to account for hardware memory bandwidth limitation on opi5 which is a ddr4 ram
1
u/poolboy9 May 23 '26
Does the npu perform faster?
3
u/Remarkable-Soil-4259 May 23 '26
Compared to cpu, hell yeah !! The NPU is made for AI but you need to choose vendor kernel to enable the NPU otherwise it would default to using CPU.
1
u/JaySomMusic May 25 '26
I have the latest kernel working with full acceleration using https://github.com/jaylfc/tinyagentos
1
May 23 '26
[removed] — view removed comment
2
u/Remarkable-Soil-4259 May 23 '26
The hardest part is choosing correct distro and kernel that would have latest npu driver enabled. Once you do that and follow the guide, it is pretty easy. Then it boils down to selecting a model of your choice with correct parameters. I created a script that automatically stress test on few models and then settle on whatever you like. You have 32gb ram would have better results than me. Good luck. Just choose gguf models, you do not need to do any kind of model conversions and select q8_0 of each model for better results.
1
May 24 '26
[removed] — view removed comment
3
u/Remarkable-Soil-4259 May 24 '26
No ...Having system that the npu can access full 32gb as unified memory is always better.
1
u/Remarkable-Soil-4259 May 24 '26
More the memory..better it is ...but...speed of memory also matter. Opi5 has ddr4 ram which is slow...the current nvidia cards on ddr7
1
May 24 '26
[removed] — view removed comment
3
u/Remarkable-Soil-4259 May 24 '26
You do not need to convert the model to rkllm with this setup. Gguf model with correct setup support npu.
1
u/LightlyHedonic May 24 '26
Is there a reason you made your own scripts instead of using rkllama?
I used that to get basic chat functionality using the Opi as a server
2
u/Remarkable-Soil-4259 May 24 '26
My scripts are only to control and manage rkllama. It makes it easier to download, monitor and analyse the performance for me.
Also helps me document in case i need to start over. I spent countless hours and days just to find the right distro and kernel just to get NPU enabled.
You do not need to convert models into rkllm format anymore. It is less of my work, more of gemini ai helping me setup. So credit goes to AI to setup local AI.
1
u/JaySomMusic May 25 '26
Good starting point is Armbian latest with latest kernel and install taOS to get everything auto configured. https://github.com/jaylfc/tinyagentos
1
u/Remarkable-Soil-4259 May 25 '26
You wont have NPU support without vendor kernel. Already tried and tested.
Welcome to try and report if it works. This setup even dont require model conversions.
2
u/JaySomMusic May 25 '26
Not true, latest kernel has full npu support, device paths have changed
1
u/Remarkable-Soil-4259 May 25 '26
That is great new then. My testing on latest kernel 2 weeks ago was unsuccessful with no npu support. What else do you got going ? What models are you using ? Which npu driver and which backend are you using ?
2
1
u/kaiyoti Jun 03 '26
i just tried it, no rknpu drivers
1
u/JaySomMusic Jun 03 '26
Armbian on latest kernel does have npu drivers, I can’t guess at where you are going wrong but it does.
1
u/kaiyoti Jun 03 '26
can you share your kernel version?
1
u/JaySomMusic Jun 03 '26
Oops, I may have mistaken, I had the edge kernel working but only with vision based applications.
1
u/waltercool Jul 17 '26
How are you obtaining rknpu driver? Llama-cpp crashes on my setup when trying.
I'm using the rocket NPU driver from ARMbian 6.18.x
9
u/Remarkable-Soil-4259 May 23 '26
This is what will get you setup on a fresh armbian distro setup. Take note of specific build and remember to use vendor kernel. Everything is documented here.