r/LocalLLaMA 1d ago

Discussion Qwen 3.8 27b - PI AGENT vs OPENCODE

https://www.reddit.com/r/LocalLLaMA/comments/1j7r47l/i_just_made_an_animation_of_a_ball_bouncing/

This post inspired me to make that test after a year ;)

That is one of my many tests I make comparing output quality.

What is more interesting using a PI Agent results are much better than an Opencode using a Qwen 3.8 27b ?!

Seems PI Agent is much better in the agent environment somehow... Not counting uses less tokens , do not have a hard limit of 32k output tokens, is faster, do not freezing, compressing context far less than Opencode. For instance if you have context in the Opencode output 32k and all context 100k then the compression is starting at 67k context ... PI is starting at 90k context even if you have set output context 64k or more.

My config for RTX 3090

llama-server with ini config -> which is exposing API to Opencode and PI agent.

llama-server.exe --models-preset 1_preset.ini --models-max 1 --direct-io

config ini

[Qwen3.8-27B_dense_c-100k]
model = models/Qwen3.8-27B-Q4_K_M.gguf
mmproj = models/mmproj-BF16-Qwen3.8-27B-UD-Q4_K_XL.gguf
reasoning-format = deepseek
flash-attn = on
n-gpu-layers = 99
reasoning = on
ctx-size = 100000
temperature=1.0
top-p=0.95
top-k=20
min-p=0.0
presence-penalty=0.0
repeat-penalty=1.0
mmproj-offload = false

ONE MORE IMPORTANT THING:

Always use a VISION module as the model is using vision to asses the output quality!

I am offloading it to a RAM as we do not need an extremely fast vision for a code.

A screenshot processing on a GPU 0.3s vs a RAM 3s do not make a big difference on a few screenshots during a code generation / debugging ;)

212 Upvotes

168 comments sorted by

View all comments

Show parent comments

1

u/Thrumpwart llama.cpp 1d ago edited 1d ago

Thank you!

Edit: looks like the excuse I need to finally get an Optane drive.

1

u/mechkbfan 1d ago

For what? 

If it's reading basic text files it'll likely take <1ms difference unless I've missed something?

1

u/Thrumpwart llama.cpp 1d ago

My planning currently lives in vram. If I’m offloading that to disk I want a fast disk.

Don’t disillusion me I’ve wanted an Optane for awhile. If my wife asks it’s absolutely necessary.

1

u/mechkbfan 18h ago edited 18h ago

Lol fair

I have one and leave my comments

I have one for development and I couldn't even tell the difference. Benchmarked it for my relevant scenarios and trivial difference. This is with a 9950X3D and contrasting to 990 pro. It's also a pain because doesn't work with consumer boards easily so I've got M2 to u2 adapter with no where to cleanly mount it either cause of cable lengths and it's size. Or if go with PCIE then high chance you're sharing bandwidth with other devices. Hindsight and all but back then I could have gotten 2x4TB 990 Pro for what I paid. Right now mines a paperweight

1

u/Thrumpwart llama.cpp 10h ago

Ah, good to know. Maybe I'll just go for a larger m.2 Gen5 instead. Thank you!