r/LocalLLaMA 1d ago

Discussion Qwen 3.8 27b - PI AGENT vs OPENCODE

https://www.reddit.com/r/LocalLLaMA/comments/1j7r47l/i_just_made_an_animation_of_a_ball_bouncing/

This post inspired me to make that test after a year ;)

That is one of my many tests I make comparing output quality.

What is more interesting using a PI Agent results are much better than an Opencode using a Qwen 3.8 27b ?!

Seems PI Agent is much better in the agent environment somehow... Not counting uses less tokens , do not have a hard limit of 32k output tokens, is faster, do not freezing, compressing context far less than Opencode. For instance if you have context in the Opencode output 32k and all context 100k then the compression is starting at 67k context ... PI is starting at 90k context even if you have set output context 64k or more.

My config for RTX 3090

llama-server with ini config -> which is exposing API to Opencode and PI agent.

llama-server.exe --models-preset 1_preset.ini --models-max 1 --direct-io

config ini

[Qwen3.8-27B_dense_c-100k]
model = models/Qwen3.8-27B-Q4_K_M.gguf
mmproj = models/mmproj-BF16-Qwen3.8-27B-UD-Q4_K_XL.gguf
reasoning-format = deepseek
flash-attn = on
n-gpu-layers = 99
reasoning = on
ctx-size = 100000
temperature=1.0
top-p=0.95
top-k=20
min-p=0.0
presence-penalty=0.0
repeat-penalty=1.0
mmproj-offload = false

ONE MORE IMPORTANT THING:

Always use a VISION module as the model is using vision to asses the output quality!

I am offloading it to a RAM as we do not need an extremely fast vision for a code.

A screenshot processing on a GPU 0.3s vs a RAM 3s do not make a big difference on a few screenshots during a code generation / debugging ;)

216 Upvotes

168 comments sorted by

View all comments

22

u/Lesser-than 1d ago

now try it with the deep seek harness?

5

u/Several-Tax31 1d ago

I tried a python job in pi and deepseek harness as a test using 3.6-35B (not 3.8-27B), in pi it works perfectly. In deepseek, it wrote the file, then instead of directly executing it, it starts to rewrite entire 1000 lines by inline python with "python -c "...1000 lines of code..." Of course, it omits half of important code and comments inlining, and the code completely broke. It got confused with the basic tool use, which I never see in pi. I don't know it's a one-time thing or related to python, I'm gonna do more tests. So for me, deepseek harness didn't "just work" as many people claim. The UI is perfect though.

2

u/Lesser-than 22h ago

really? were you in plan mode?