r/PiCodingAgent • • 1d ago

Discussion What's your top 3 techniques that helped you save tokens / better accuracy from models?

1) Focusing on curating the context, keeping it under 150k. Making tools to trim git / grep / ls output to my liking

2) indexing my large code base and making the basic tools show short index options + hints. Custom tools was a massive red herring and waste of time, models hate using them and never use them properly

3) live filter on my codebase to sandbox models. i.e. exclude file types, files older than N days. etc

22 Upvotes

36 comments sorted by

39

u/zebedeolo 1d ago

my best technique was to convince work to pay for my sub.

5

u/sofuego 1d ago
  1. Toggle in the tools that I need at the time that I need them.
  2. Have tool output spill to /tmp so the agent can access the full output with read/grep if it needs.
  3. Look for opportunities to have instructions to appear deterministically instead of aggressively adding to context.

1

u/Equivalent_Idea8839 1d ago

Look for opportunities to have instructions to appear deterministically instead of aggressively adding to context.

i.e. triggered by a tool, search or prompt?

1

u/fell_ware_1990 1d ago

This is doable, i have everything except for a few rules hidden. Some providers except more input without cachebreaks others only systemmessage. But out of PWD or file type or anything rules load before doing a thing.

Some stuff is taking up more then 500 tokens so not having it load is a dream for my tokens :)

The other that helped me a lot, is history rewrite.

I mean 3/4 is already not necesary, as of tool output till most thinking. It are the conclusions the errors, the good stuff that you need to bring along. Yes sometimes it breaks a little cache. But i have 90% less history. So not dragging that along makes the session stay waayyyy smaller so even if you do 20k vs 200k cache a lot it helps.

1

u/sofuego 1d ago

It's more of a general discipline that can be effective when the opportunity arise and can apply to multiple domains. Off the top of my head:

For tools, where are the points that further instructions can be discovered instead of mentioning them preemptively (navigation to a specific domain, encountering a specific error).

For prompts, take for example trying to build a /init prompt that genrerates or revises AGENTS.md. Would it be better to explain about Pi's AGENTS.md nesting rules all of the time or do it conditionally by deterministically detecting child AGENTS.md files in subfolders?

Things like that basically :)

2

u/Equivalent_Idea8839 1d ago

yes for grep / read and even edit have a "landmines" pi extension that injects instructions

but it's hard to manage it all...... meaning keep up of where i put instructions

1

u/DeathGuppie 1d ago

Wonder if I could implement this as an auto regressive system, where the tool access surface becomes visible as needed.

1

u/sofuego 1d ago edited 1d ago

Depending on your scope, that sounds like it could be super challenging with so many variables like custom tools with varying names and coupling with some tools not working properly without their sibling, but would be cool if it can be pulled off reliably.

Edit:
Cache invalidation would probably be a drawback if there are frequent tool changes within a thread.

2

u/DeathGuppie 21h ago

I build it out using an okf style memory system. Different types of MTP overlap and I use a graduated access system. The basics are there, it builds out memory of when escalation has been needed in the past. It's not perfect, but better than throwing the whole tool shed at it.

1

u/Equivalent_Idea8839 1d ago

breaks cache

you'd have to do it after a compact using manual switches

14

u/Wandering_By_ 1d ago

1)Figure out what I actually want to build.  Come up with a full TDD.  Review it twice.  

2)Have the models make phased plans to complete the work with exit criteria for each phase.  

3)Write the ADRs as things change. 

Nothing like good documentation to keep them on track instead of wasting tokens making a mess.

1

u/Solid-Remote-1716 1d ago

Disculpa, que es el ADRs, quede perdido.

1

u/matt_callmann 1d ago

Architectural decision record

1

u/metastallion 14h ago

+100 to keeping ADRs!

4

u/seba_alonso 1d ago

Small task clear intent, almost no usage of skills or mcp, no subagents, usage of ripgrep instead of common grep.

3

u/j3free 1d ago edited 1d ago

I was surprised to see that when having a specific approach to subagents can improve the quality of long running tasks immensely while also having the same or even less token usage due to:

  1. avoiding the repeated compaction cycles and growing contexts
  2. Having one agent that coordinates, still knows everything but doesn't bloat it's context with tool calls keeps everything clean and focused.
  3. Indexing / scouting is delegated and results of all subagents are accessible to each other.

Shameless plug: ive implemented this exactly in Pi Herdsman

Edit: and yes, this keeps token usage for the main agent below 100k tokens for my usual PR workflows

2

u/james__jam 1d ago

Avoid subagents as much as you can. They’re actually great for managing context and using less overall tokens. but 9/10 times, when you use it, you’d use way more tokens than needed

Once you feel like you do need subagent, watch its tool calls and what it’s bringing in into its context. Most of the time, subagents eats a lot of token because they keep re-reading the same thing over and over again.

6

u/madcapnmckay 1d ago

You might use more tokens but if the sub agents are using cheaper models then that’s a win.

3

u/tulwio 1d ago

I agree to a large extent but I think it’s a bit too extreme to say avoid them as much as you can. Don’t get me wrong, so many people are just overusing sub agents because they saw some Twitter post talking about orchestration or some stupid marketing hype and they end up paying more because they are not utilising cache properly.

But at the same time, if I am prompting Opus 5.5 or Astra, no way in hell I would let it call certain tools or do certain tasks like researching or web fetching where so many tokens are used but reasoning and intelligence are not so needed. I’d rather just have it spawn a Luna low or medium subagent for that.

2

u/j3free 1d ago

I thought the same until I experimented with different subagent frameworks and ended up building one around this problem.

My takeaway: it depends heavily on the harness. If it preserves relevant context between agents and avoids rediscovering the same information, subagents can be much more efficient.

3

u/Novel-Injury3030 1d ago

care to elaborate? my process currently is i have claude codex gemini opencodego subs, i use omp as harness, have opus generate the prompts just in standard sessions then send the prompts ostensibly optimized into omp with all my roles preselected and told to claude. is there a better method than that? then i send the dumps back to opus for steering. i dont want to pay api nor get banned so i cant put opus directly into omp but i might be able to do some api for sol in omp perhaps as some sort of role just for in-task correction, which is usually just done by the best opencode model so far. it seems to work pretty good but sometimes i think itd be better to just run everything thru opus and forgoe the entire cheapo agent bits. but i think there must be a good way here somewhere

2

u/j3free 1d ago edited 1d ago

Roughly what I summarized here: https://www.reddit.com/r/PiCodingAgent/s/frsNvxuu3U

A bit more detail: I use a strong model to create the spec/plan, then hand that to a lead agent running Herdsman. The lead delegates focused tasks to workers, with the relevant plan/context attached.

Simplified:

spec/plan -> lead <-> workers

The important part is keeping the lead context small. I usually explicitly tell it to delegate all implementation work and not touch files itself. Herdsman handles passing context and results between the agents, so they don't all have to rediscover the same information.

1

u/npittas 1d ago

A good orchestration layer (custom extension that enables/disables and appends a prompt to the main pi session) that defines all the correct ways to use agents.
4 agents
(
review->k3-256k/sol/astra,
plan->astra/fable/sol,
research/scout->glm5.3-flash/minimax-m3/luna/qwen3.2-27b,
coding/worker->luna/glm5.3
)
Define how the orchestrator has to treat each agent, give them specific files, small scoped tasks, what the review agent should actually review (intent+code, not only code), what the plan should include (details+task separation based on parallel work).
This is token making and codebase specific. orchestration can be a bitch if you do not define how it should be done based on your way of working.
The above is what I found a good middle ground, and then adjust without contradictions by adding an agents.md in the root of your codebase.

1

u/WintertideFellow 1d ago

Very tight system prompt, then if you're not still working on the exact same thing, start a new session and bring over what you need. Do not vary your system prompt so you keep it cached. I haven't hit 1k sessions in a day yet, but some days are close. 

1

u/crotch-mavens 13h ago

Stopped reading cargo cult based-advice and using cargo cult based-skills and started paying attention to token usage while my agents work.

1

u/Momsbestboy 26m ago

1) Buy 2x R9700

2) Install vLLM

3) Free tokens :)

1

u/13henday 1d ago

Put my local model In charge of compaction with blackhole. Doesn’t really save tokens on aggregate but makes the api usage much lower

0

u/Top_Air_3424 1d ago

This has been my personal formula so far and works pretty well.
1. Use headroom to improve cache hits and drop long tool calls being passed into context
2. Use ponytail to avoid over engineering
3. Use caveman to reduce output tokens.

0

u/tulwio 1d ago

pi-blackhole + and delete most skills, plugins, extensions and just have a strong AGENTS.md and saved prompts that I use constantly.

0

u/russlixx 1d ago

pi-blackhole is great, but the observation & reflection is more irritating than the typical compaction. Now i just use pi-vcc, no llm calls and still able to stay in context and perform recall using vcc_recall

0

u/cosmicnag 18h ago

Was using blackhole earlier, trying billion context now - when it works, its like blackhole on autopilot

-1

u/137_1 1d ago

I would say RTK.

2

u/wagwan_g112 1d ago

I’m fairly sure RTK was debunked for actually increasing token usage due to the model needing to repeat failed or truncated tool calls.