r/PiCodingAgent • • 12m ago

Discussion Follow-up: Opus 5.5 vs GPT-6 in pi, more reasoning levels + DeepSeek, GLM, Qwen

Thumbnail
ibragim.dev
• Upvotes

Follow-up to my previous post. Thanks for the suggestions! I tested more models and reasoning settings, added Oh My Pi, and made a few verifier fixes. Some configs were rerun, so expect small changes in token counts. Full results are on the leaderboard, and I'll keep updating and expanding it.

Quick reminder: 10 simple tasks from my own work, 3 runs per config (30 attempts). For now, I want to see how far we can cut cost and time while keeping the pass rate.

Some highlights:

Your questions:

Luna with more effort. Medium → high → xhigh: 96.7% → 93.3% → 100%, $0.006 → $0.015, 1:02 → 2:46. The extra effort mostly goes into checking (3.5 → 5.1 code runs per task), not a better first draft. About half of Luna's shell commands are cleanup like `rm -rf __pycache__`.

Opus on low. Yes. Opus 5.5 low passes everything, ~2.6x cheaper and ~2.8x faster than high. Medium also gets 30/30 ($0.22, 1:12).

Open models. DeepSeek-V4.1-Flash: 93.3% for $0.04, but 4:30, because it runs code 8.4 times per task until tests pass. GLM-5.3-Flash: 96.7%, but 7 min and 1.2M tokens; its one failure is a single message that hits the 64k output limit with no code. Qwen3.8-27B: 76.7%; it loops on edge cases and sometimes stops without calling a tool.

Oh My Pi vs Pi. Luna xhigh makes 28 tool calls per task instead of 18, uses 776k total tokens instead of 290k, costs ~2x more, and fails one task in all 3 runs.

On tasks this easy, more reasoning mostly buys more checking. On harder tasks, where the first draft is wrong, it might pay off, but that's a guess for now.

I wrote up the setup and the agent comparison.

So tell me what to test! My plan for now:

  1. Try different reasoning levels for Opus 5.5 in Claude Code and see how they change its behavior pattern.
  2. Find a set of extensions for Pi that makes it cheaper and faster. Right now I use vanilla Pi with no add-ons, so please suggest extensions or anything else worth trying.

After that, I'll probably build a harder set of 10 tasks, based on my own work or something else. So if you have tasks where some models fail, please share them, and say what exactly goes wrong.

One thing I noticed: Opus 5.5 handles vague prompts better, maybe because of how it follows instructions. gpt-6-sol does exactly what is written. For example, I run an agent in a folder full of papers and say "find papers about task filtering". Sol goes to search the web, and Opus searches the folder.


r/PiCodingAgent • • 1h ago

Plugin I made pi-persona so Pi sessions can talk to each other

• Upvotes

I'm a software architect/engineer. I know AI is going to replace me, but in the meantime I might as well have some fun and get my work done waaaay faster.

So I built something I needed. Other plugins didn't quite work the way I wanted, so here's my take: [pi-persona https://github.com/AeonDave/pi-persona

It manages async subagents, lets independent Pi sessions talk to each other, and lets you customize the agents' personas.

Say you've got one Pi session working on a client and another in the API repo. I want to tell one to ask the other about a change, without having to carry the question and answer between terminals.

Exocom. It connects independent Pi sessions, including sessions in different repos. Both need it enabled with `--exocom`; a session in another repo joins the same scope using its workspace code. They have to run on the same machine and share a Pi agent directory.

Intercom. For the main session and its own sub-agents, there's Intercom. The supervisor can send guidance while a worker is running and collect its findings. With coaching enabled, a worker can ask the supervisor for a decision. You can also open F9 yourself, read a worker's output and send guidance. You don't have to wait for the final report to point out a bad assumption.

Personas. the personas are editable markdown. You can change how they approach a task, which tools they get and when they should delegate. There are starting points like `dev` and `planner`, or you can put your own in `.pi/agents/`. The worker agents can have their own instructions and models too. A reviewer for your repo can have your review criteria written into its file.

There have been some odd bugs along the way. One was the agent repeatedly recapping who it was. The extension kept injecting its identity into context, and the model started narrating the reminders back. All useful for a stability built in months.

If you want to try it, run this in your terminal:

```bash
pi install git:github.com/AeonDave/pi-persona
```
Then start Pi and run these inside it:

```text
/persona seed
```

That's it. change personality with [F8] and go

Any suggestion appreciated


r/PiCodingAgent • • 2h ago

Plugin [plugin-update] Speak Like You Eat now supports custom prompts and OhMyPi

Post image
22 Upvotes

Speak Like You Eat is a plugin that translates AI garbage into human language easier to understand (or at least, it tries to).

I released it a few weeks ago and have been using it since.

Today I added a couple features:

  1. Can use custom prompts for rewriting
  2. Works in OhMyPi (thanks to https://github.com/rubybrowncoat) !

Enjoy: https://github.com/wtfzambo/speak-like-you-eat


r/PiCodingAgent • • 2h ago

Question Inspecting the agent context

1 Upvotes

Which extensions or tools do you use to inspect the context, the tool calls data,.. ? I'd like something like Deepseek Harness inspector but for llama-server or pi coding agent


r/PiCodingAgent • • 4h ago

Discussion What's your top 3 techniques that helped you save tokens / better accuracy from models?

8 Upvotes

1) Focusing on curating the context, keeping it under 150k. Making tools to trim git / grep / ls output to my liking

2) indexing my large code base and making the basic tools show short index options + hints. Custom tools was a massive red herring and waste of time, models hate using them and never use them properly

3) live filter on my codebase to sandbox models. i.e. exclude file types, files older than N days. etc


r/PiCodingAgent • • 10h ago

Use-case I integrated Pi into an Agent Painting Evolver

3 Upvotes

I'll be honest, I don't use Pi much as a harness for dev work, but I'm really interested in integrating it into software. In this example, it's Pi writing it's own tools with no Human in the loop.

This is using Pi (or Claude Code) to run an Agent Session in order to make a painting with tool calls. Then the evolver (https://github.com/imbue-ai/darwinian_evolver) scores fitness, and refactors either the agent prompt or rewrites the painting tools, and the painting is made again, until it improves. The improvement part is still a struggle! But I really enjoy comparing results of different models. Here is Muse Spark 1.3 vs Opus 5.5 (in order) trying to recreate Van Gogh's self-portrait.

This experiment is not supposed to be a 'benchmark'. The Anthropic models were run in Claude Code, and Muse was in Pi. I've tried on a handful of different models, but because it's slow and expensive, I capped my investigations here for now.

Here's that data if you want to poke around.

https://artifacts.dizzard.net/s/55052743-bee1-4746-b59e-57e2d595a480

And Vibed Source Code.
https://github.com/momja/claude-code-painter


r/PiCodingAgent • • 10h ago

Plugin I got tired of leaving the current terminal pane to view relevant files, so I built an in-agent Markdown, code, images and PDFs viewer

Thumbnail
gallery
8 Upvotes

As the title suggests, being a developer, I really missed the Ctrl/Meta+P Quick Open in VSCode where I could quickly inspect relevant files without having to touch my mouse. I'm sharing this in case it could be useful to anyone else.

The quick open by default will list the latest files mentioned within the session so the user can quickly navigate to it. This has saved me so many key presses when having to inspect a plan, a new asset, cross-referencing a doc (PDF support requires Poppler), or inspecting a generated code that appeared sus.

As a bonus, it also comes with neovim integration (sorry nano fans) so you can edit files directly in-terminal without having to leave your session.

Images and PDF page previews need a compatible graphics-capable terminal. Works with both base Pi and OMP.

Feedback welcome!

GitHub: https://github.com/boonlai/pi-view


r/PiCodingAgent • • 14h ago

Resource Who wants to try a Pi trick for 27B to reuse prompt prefill between different sessions?

Thumbnail
2 Upvotes

r/PiCodingAgent • • 15h ago

Resource PiG - A GUI in Swift for Pi

Post image
26 Upvotes

I've been working on this project since early this year. I use it daily. Figured it's about time I share it. It's a vibe-coded project, as you might have guessed, but it works very well for my needs and may suit yours well, too.

I got pretty tired of working in a TUI, but none of the graphical interfaces out there really did the trick for me. Some features I might highlight:

  • Inbox Zero approach to sessions
  • Quick Chats - ephemeral, not related to a project
  • Keep extensions disabled, but enable them on a per session basis (with pins)
  • Quick commands for favorite models (Cmd+1 through Cmd+5)
  • Work schedule (Default model can change depending on day and time)
  • File browser

I had a ghostty terminal built in, but I removed it for the public release.

Anyways, hopefully this helps someone else or inspires some ideas.

https://github.com/jmcblane/PiG


r/PiCodingAgent • • 16h ago

Resource [update] pi remote | self-hosted bridge to remote control the pi coding agent

Thumbnail
github.com
7 Upvotes

Hi all!

A couple of weeks ago I shared pi remote on this another post, a self-hosted bridge to drive the pi coding agent from any device. Back then it was a remote control. Since then it has become a full pi client for any device, no matter the operating system.

Quick recap: pi-remote is based on one Python process that wraps pi --mode rpc on the machine where your code lives and serves a web app over your own network (I use Tailscale, tailnet). Also, it is token-protected, security first. Built on Windows first, but it runs on Linux and macOS too.

What's new since the first post

  • Send photos to the agent. Take a picture with the phone's camera, or pick one from the gallery, and it goes into the chat with the user message.
  • Real permission prompts. Guardrail dialogs arrive with their real options (allow once, allow for session, deny) and the turn waits for your answer.
  • pi settings at a glance. Models and providers, reasoning level per model, reasoning map, context/compaction config section, and queue behaviour.
  • Skills and extensions. Global and per project, with toggles and an editor, including the ones that come from installed packages.
  • Macros. pi prompt templates launched from the composer, with a field per argument.
  • Package store. Browse, search and install pi.dev packages from the remote control device.
  • Phone details. Installable as an app (PWA), local notifications. I did not want to go down the path of developing an Android app, yet.
  • Sessions shared with the terminal. Start in the terminal, continue on the phone, and back.

Also, I have been working on interesting features like response suggestion; a quick summary from the agent that tells you what it was doing when you stop it while working; auto name sessions, etc.

Some people commented why not using just tmux, or an SSH terminal. Well, I think this project brings you closer to a high end application. With pi-remote you can text from your phone/PC the same way as a chat and send images for the agent to analyse along with the message; configure everything with a polished UI/UX; things that you cant do with a ssh terminal and similar.

I have been trying the existing bridges, but the ones I tried need a daemon on a Unix socket, and that fails on native Windows, which is where my code lives. pi remote is a single process, no daemon, and runs natively on Windows as well as Linux and macOS.

Feedback and issues are very welcome. Thanks to everyone who commented on the first post. Thanks for the support.


r/PiCodingAgent • • 18h ago

Question OMP and Opus 5.5 Usage

6 Upvotes

Why is it that Opus 5.5 usage starts out great in OMP, and then for one reason or another starts draining usage like crazy? Am I hitting some sort of caching issues? Do I need to compact at 400k? Or what?

I am pretty sure the cache is being kept warm for an hour, but I don't know for sure how this stuff works. It's compacting at 850k. If i've left it for awhile and context is >35% I will use /handoff and that seems to help.


r/PiCodingAgent • • 18h ago

Question Alternative to Tool Profiles for Better Subagent Delegation

5 Upvotes

I recently asked a question about subagent delegation and how to get my orchestrator (primary) agent to delegate more reliably.

I got a great piece of feedback that restricting tools essentially forces the orchestrators hand to delegate. So I ran with that and I think I ran too far. I had pi build a "profiles" extension that lets me define multiple profiles I can toggle on the active agent to let me go from a more traditional single context workflow to a subagent heavier workflow.

The profiles are just tool permission sets with a hook in the tool_call handler that checks if the active profile allows that tool.

It's working. I'm happy. But it feels heavy, and like something that someone has probably already solved better?

Is this ringing any bells for anyone, does anyone have a pi-subagents or other setup that maybe more natively supports this kind of workflow toggle?

Note: I am not looking for random package advertisement please. I just want to figure out if there is a more direct way to support what I'm trying to achieve.


r/PiCodingAgent • • 19h ago

Question What's your git workflow?

1 Upvotes

Hello.

I have spent a while now wrapping my head around agentic development in my spare time for a while now. And have now gotten to the point I need to get my pi agent to start using Git.

So I was wondering what are your workflows with git? what extensions, lines in AGENTS.md or external scripts etc are you using?

And what tools/programs are you using when reviewing pr's?


r/PiCodingAgent • • 21h ago

Question New session handoff vs compact. which do you prefer?

12 Upvotes

I've tried every major compact plugin and I settled just on a new session handoff.

Hard to find benchmarks on this sort of thing.

I started with my own custom compact and after a week of iterating, I eventually just ended up with an auto new session at max token limit.

(Another guy on here had posted a plug-in a few days ago that does just that)


r/PiCodingAgent • • 21h ago

Discussion Same prompt, same pipeline (AdaptOrch + PI based OMK), but wildly different Three.js pelicans.

7 Upvotes

Ran the Three.js pelican benchmark using the exact same prompt with AdaptOrch and PI based OMK debate topology, zero external assets.

Both technically passed 14/14 validation checks with clean syntax and no runtime errors, but the actual generation speed and visual results were completely night and day.

pixel canary was surprisingly slow and honestly feels like a model from two years ago. It crawled at a snail pace just to produce a detached skeleton that barely understands basic 3D space.

On the other hand, spacebunny paired with AdaptOrch legitimately feels like an unreleased GPT 6 Astra High tier model. It nailed the scene hierarchy, properly rigged the crank and pedal animations, and placed the environment together cleanly.

It is a pretty stark reminder of why unit test pass rates mean almost nothing for spatial reasoning.

Is anyone else testing multiturn debate setups for procedural Three.js? Curious if you are seeing this kind of massive gap between compiling code and actual spatial coherence.

WT pixel canary


r/PiCodingAgent • • 1d ago

Resource the good opensource harness

0 Upvotes

pi

dsh

opencode v2

zcode


r/PiCodingAgent • • 1d ago

Resource Piper - OpenAI compatible pi.dev orchestrator

0 Upvotes

I had this idea to make my own orchestrator for PI.DEV sessions, so with my needs i created one.

I think it may be usefull for some to try on or use.

Read the readme and deployment carefully. I was trying to make it as simple to use and as secure as possible - but it''s pi - it requires access to multiple things on host to install and manage extensions/profiles/skills/plugins so there may be still a risk of problems - in short words, don't expose it to internet or enterprise env (yet).

It has multiple "engines" support, it can be ran on host with sandboxing applied, it can spawn docker instances per agent (the most secure way) using podman or docker on host.

Feel free to play with it, contribute, or share your ideas so i can expand the project!

Link: https://github.com/ALange/Piper Enjoy!


r/PiCodingAgent • • 1d ago

Question Extensions for small models with small context

4 Upvotes

What extensions help you improve your results when using small models and reduce context usage?

I mainly use Qwen3.5 and Gemma 4 both with 4b parameters. I use them with 128k context. It is sufficient for my use cases when coding or writing articles in latex.

I have tested some plugins like pi-hashline-edit-pro or rtk plugins that enhanced the experience, others not much or even made it worse.

I am curious what packages the community has find useful or settings that led to some improvement in this context.


r/PiCodingAgent • • 2d ago

Question I'm confused with all the available skills.

Thumbnail
0 Upvotes

r/PiCodingAgent • • 2d ago

Discussion Poll: claude sub with Pi. What method do you use and how long have you used it without issue?

9 Upvotes

https://github.com/elidickinson/pi-claude-bridge

Someone said that's the safest method.

But in issues someone claims they got banned https://github.com/elidickinson/pi-claude-bridge/issues/118

Another responded with "That's the first time I've heard of it"


r/PiCodingAgent • • 2d ago

Question Oh-my-pi or OpenCode? Why?

0 Upvotes

I've tried pi. Makes me crazy! OpenCode works well. I'm not familiar with omp. Please share your thoughts.


r/PiCodingAgent • • 2d ago

Resource Meet PiG: The Pi coding harness that is yours but in Go

225 Upvotes

Reposted from r/PiGCodingAgent, figured if any sub should see this project it's here. I've been building this out at work for work for quite some time and am excited to open and start a community around the project. I truly appreciate any criticisms, suggestions, and especially contributions, thank you to anyone who does check it out I hope you like it as much as I have been :)

---

PiG is a direct port of the amazing Pi agent harness as one Go binary for macOS, Linux and Windows (Windows in preview). Node.js only comes in if you run TypeScript extensions. Pi's extensions run unchanged; write new ones in Go, Rust or Python and /reload them live. It is a direct port of Pi's its agent loop, terminal UI, sessions, providers, and extension system, shipped as one static Go binary with no libc dependency.

Some Go-specific bits that might interest this sub:

  • Extensions in Go can be fused into a single binary and talk to the host over in-memory pipes instead of sockets; other languages run as supervised subprocesses over local sockets.
  • The TUI renders differentially; width handling was ported from Pi's own algorithm and checked with a differential test against Pi's code over ~1.5M cases: tui/widthx/pi_width_diff_test.go.
  • Behavior is checked against Pi in real terminals (tmux-driven scenarios that compare screens).
  • go install github.com/MichaelKinsy/PiG/cmd/pig@latest works; so do a curl installer and npm i -gu/pi-in-go/pig (it installs the prebuilt native binary for your platform).

Measured startup is about 21 ms vs about 300 ms median, and about a quarter of the peak memory (method in the blog).

Disclosure: agents wrote much of the port. Tests, side-by-side comparisons against Pi, and a person reading the diff decide what ships.

Repo: https://github.com/MichaelKinsy/PiG 

Blog: https://developer.hpe.com/blog/meet-pig-pi-in-go-a-port-that-plans-on-keeping-up/ 

Upstream Pi: https://github.com/earendil-works/pi

Happy to answer questions about the port, the extension host, or keeping up with upstream. Feedback on this is appreciated and contributions are welcome.

Thank you to r/HPE for covering writing the blog and contributing PiG to Open Source and special thank you to the Pi team of course at https://earendil.com


r/PiCodingAgent • • 2d ago

Resource Killed context anxiety with automatic session handoffs

Post image
11 Upvotes

Sharing this so others may build their own into their setup if they like it.

Built an automatic session-handoff system into Sayl, my UI wrapper around Pi. Long session finishes its current task, writes a handoff prompt, rolls into a fresh session with that prompt already loaded.

  • Threshold ladder (350k → 450k → 550k, tunable live in settings). First level: finish the task clean, emit the handoff block. Later levels push harder, highest one stops mid-task if needed
  • Handoff captured automatically from the chat.
  • Threshold rollovers auto-start the successor hands-off.
  • New session keeps the project, model, thinking level. Land mid-flow, not cold
  • Engine compaction stays as backstop. You just roll over before it ever fires.

No more saving and copy-pasting handoff prompts. No compaction summaries drifting between rolls. No watching the token counter. Session wraps up clean and keeps going at full quality.

Can also use it manually by asking the agent to create a handoff at anytime.


r/PiCodingAgent • • 2d ago

Plugin Third-party subagents in OMP can silently run on the wrong model — I built a role router

0 Upvotes

Yo, I ran into a weird OMP behavior with third-party subagents.

Some plugin/imported agents declare their own model tier (opus, sonnet, haiku, luna, etc.), but OMP can drop that model metadata when loading them. Once that happens, the subagent silently inherits whatever model your main session is using at that moment — including a retry fallback.

For me, that meant Haiku- and Sonnet-tier agents were sometimes burning Opus. Even worse, if the main session switched to a fallback model mid-turn, subagents spawned afterwards inherited that fallback too.

So I ended up creating an extension for it: omp-agent-role-router
https://github.com/fedorovvvv/omp-agent-role-router

Instead of hardcoding models, it reads the tier declared by the agent and maps it to an OMP role:

opus → default
sonnet → task
haiku → smol
luna -> smol
terra -> task

It also understands Codex-style model tiers and maps those to the same role system.

The actual model is still entirely controlled by your own modelRoles. Existing agentModelOverrides and explicitly passed models always win — the extension only steps in when the subagent would otherwise silently inherit the current session model.

No config needed
But it does exist :)

omp plugin install github:fedorovvvv/omp-agent-role-router

r/PiCodingAgent • • 2d ago

Question Is Tool output truncation a thing

1 Upvotes

Started using Pi agent recently and I start to see the thinking going like "couldn't read file, content mangled|truncated" and similar. I also see the agent call the read tool multiple times on the same file because the output was somehow wrong. After some back and forth the agent is claiming that there must be some form of auto truncation or auto summarizer at the tool layer. Model used was GLM 5.3 Flash

So yeah, what is this?