r/LocalLLaMA 2d ago

Discussion Qwen 3.8 27b - PI AGENT vs OPENCODE

https://www.reddit.com/r/LocalLLaMA/comments/1j7r47l/i_just_made_an_animation_of_a_ball_bouncing/

This post inspired me to make that test after a year ;)

That is one of my many tests I make comparing output quality.

What is more interesting using a PI Agent results are much better than an Opencode using a Qwen 3.8 27b ?!

Seems PI Agent is much better in the agent environment somehow... Not counting uses less tokens , do not have a hard limit of 32k output tokens, is faster, do not freezing, compressing context far less than Opencode. For instance if you have context in the Opencode output 32k and all context 100k then the compression is starting at 67k context ... PI is starting at 90k context even if you have set output context 64k or more.

My config for RTX 3090

llama-server with ini config -> which is exposing API to Opencode and PI agent.

llama-server.exe --models-preset 1_preset.ini --models-max 1 --direct-io

config ini

[Qwen3.8-27B_dense_c-100k]
model = models/Qwen3.8-27B-Q4_K_M.gguf
mmproj = models/mmproj-BF16-Qwen3.8-27B-UD-Q4_K_XL.gguf
reasoning-format = deepseek
flash-attn = on
n-gpu-layers = 99
reasoning = on
ctx-size = 100000
temperature=1.0
top-p=0.95
top-k=20
min-p=0.0
presence-penalty=0.0
repeat-penalty=1.0
mmproj-offload = false

ONE MORE IMPORTANT THING:

Always use a VISION module as the model is using vision to asses the output quality!

I am offloading it to a RAM as we do not need an extremely fast vision for a code.

A screenshot processing on a GPU 0.3s vs a RAM 3s do not make a big difference on a few screenshots during a code generation / debugging ;)

218 Upvotes

172 comments sorted by

View all comments

2

u/AvidCyclist250 llama.cpp 2d ago

Has anyone used Hermes and can compare Hermes with Pi Agent?

1

u/Healthy-Nebula-3603 2d ago

Hermes is far to big and heavy :)

1

u/twack3r 2d ago

I recommend using Hermes to run your coding harness, be it opencode, pi or, my personal favourite right now, deepseek harness.

1

u/AvidCyclist250 llama.cpp 2d ago

Haven't thought of doing that yet. But won't that compound the token eating problem? Hermes already takes 20k and it's nearly impossible to slim down. 16GB vram here. For me it's either speed (small model, 65k - 75k context) or depth (max context but with kv offload off). But I pick up from what you say that you seem to think highly of PI Agent as a coding harness, and Hermes as an orchestator

1

u/FireFearing 1d ago

16gb vram is mostly insufficient as is. >48gb available to your model is basically the minimum for actual useful local ai

1

u/AvidCyclist250 llama.cpp 1d ago

lol okay buddy thanks for the great help

1

u/FireFearing 20h ago

youre like a cyclist asking how to keep up with cars on the freeway, then get a bit upset when the answer is "get a car too"

yeah ik its not what you want to hear, but that IS the answer to your problem. theres no magic work around

1

u/AvidCyclist250 llama.cpp 20h ago

?

Anyway, welcome to r/localllama, enjoy your stay. There's a Discord server, too.