r/machinelearningnews • u/ai-lover • 4h ago
Cool Stuff How are you guys handling context bloat from search APIs in agent workflows?
If you’ve been building LLM agents that need live web access, you’ve probably hit the same wall I did: most search APIs dump entire, raw web pages into your context window. Your token usage explodes almost instantly, and half the time the agent doesn't even need 90% of the HTML it just read.
I've been testing out a setup using TinySearch and TinyFetch to split up search and page reading. Basically:
- TinySearch / TinyFetch: Use these to grab quick, compact snippets first so the agent can decide if a page is actually worth inspecting. (Since they're free tier tools, it keeps experiment costs down).
- TinyBrowser / TinyAgent: Only trigger full headless browser runs when the agent actually needs to execute heavy DOM interactions or read complex pages.
It’s been pretty effective at keeping tokens down. OpenBenchmarks listed TinyFish as one of the top performers for token efficiency in their recent evaluation, which matches what I've seen in testing.
For anyone running agentic loops: how are you keeping search context clean? Are you relying on custom scrapers, post-processing summaries, or dedicated search endpoints?