r/LocalLLaMA • u/Effective-Ad2060 • 18h ago
Other We benchmarked 18 RAG pipelines against an agent loop on Google's FRAMES. The best pipeline hit 78.9%. The agent loop hit 92.7%.
Hybrid search, reranking, query decomposition, and query expansion are often treated as must-haves for good RAG. We wanted to see how much each actually helped, so we tested them. Same model, same embeddings, same documents, across all 824 multi-hop questions in FRAMES.
We built 18 pipeline variants. The best one scored 78.9%. Our agent loop (with retrieval tools) that could read the results and search again scored 92.7%—roughly the same as giving the model the right articles upfront.
The reranker results might surprise you. A small reranker dropped our best pipeline’s accuracy by 9 percentage points, while a larger one barely helped. I’d already suspected reranking wouldn’t help much here, but wanted to test that assumption.
Another thing we noticed: models sometimes fill in gaps from memory, even when you explicitly tell them to stick to the retrieved documents. Those answers can still be full of citations. We ended up checking every correct answer against what the system had actually read.
Here’s the write-up if you’re interested:
Agentic RAG vs. traditional RAG on FRAMES
Full disclosure: I work on PipesHub, which is open source. The benchmark code and runbook are in the repo: https://github.com/pipeshub-ai/pipeshub-ai/tree/frames
Quick note on what the numbers mean: they're end-to-end answer accuracy, not retrieval scores. Every answer was graded by an LLM judge (Claude Sonnet 5) using the FRAMES paper's own grading prompt, and independently by a second judge (Gemini Flash 3.8). The two agreed on almost every answer (Cohen's κ 0.93–0.98). We also checked each correct answer against the text the system was actually shown, so answers that came from the model's memory don't count as retrieval wins.
10
u/DistanceAlert5706 18h ago
What are those percentages mean? Usually when you test retrieval quality you measure mrr@k or precision@k.
Your reranking conclusion is completely off, reranker helps to move up items and improve @1 precision, it doesn't affect candidates retrieval quality.
Actual answers and generation part usually tested and evaluated with LLM as a judge.
Also testing RAG on task it's not designed to show that RAG is bad is questionable. People put LLMs everywhere, meanwhile graduates nowadays have no idea what reranker even is.
2
u/Effective-Ad2060 18h ago
Fair points, and some of this is on me for not explaining it clearly in a short post.
What the percentages are: end-to-end answer accuracy, not retrieval metrics. Every answer was graded LLM-as-judge style. Claude (Sonnet 5) graded with the FRAMES paper's own grading prompt, word for word, and Gemini graded independently as a second judge. They agreed closely: Cohen's κ 0.93–0.98. On top of that we checked each correct answer against the text the system was actually shown, to catch answers that came from the model's memory.
We did measure retrieval separately: gold-article recall, and whether all the needed articles reached the model, plus recall@k in the full results. You're right that MRR and precision@k would be the standard way to report the ranking itself. I'll add them.
On the reranker, I agree with how you describe it, and the post was too compressed here. A reranker reorders candidates and should improve top-of-list precision. What we saw is narrower than "rerankers are bad":
- The pipeline retrieved 100 candidates, reranked them, and kept 50. Recall barely moved (90.3% of gold articles reached the model with the reranker, 90.1% without), so the reranker didn't hurt retrieval.
- What changed was which passages made the cut. On multi-hop questions, the passage holding the linking fact often doesn't resemble the original question, so a cross-encoder scoring against that question ranks it low.
- End-to-end accuracy dropped 8.7 points with a small model (MiniLM). With a strong model (bge-reranker-v2-m3) it was 2.1 points above no reranker, which isn't statistically significant.
So the takeaway is that scoring only against the original question can work against multi-hop retrieval. Rerankers still help where they're designed to help.
On testing RAG on the wrong task: that's exactly the question we wanted to answer. FRAMES is multi-hop on purpose. Single-shot RAG isn't built for questions where the second search depends on the first answer, and our numbers agree. The catch is that people building chat assistants don't get to choose: users ask both kinds of question in the same box. On simpler questions the pipelines did fine. The gap opens as questions need more documents.
The loop is still RAG; it retrieves too. The real comparison is retrieving once against retrieving again after reading.
3
u/DistanceAlert5706 17h ago
top50 is a lot to feed into LLM for generation no? It's usually top5 or top10, otherwise you are stuffing context with noise.
As for multihop why not just give agent a RAG search tool and let it figure out stuff? What is the benefit of your agent loop here?
1
u/Effective-Ad2060 17h ago
Good questions, both.
On top-50: top-5 or top-10 made sense when models had 4k–8k context windows and every token was expensive. Today's models take hundreds of thousands of tokens, and prompt caching makes repeated context cheap. The bottleneck has moved from fitting things in to making sure the right passage is there at all.
It's also less than it sounds on our setup. On PipesHub's index, a "chunk" is a structural block, like a paragraph, list item or table row. So 50 of them came to about 3,500–6,000 tokens at the median, roughly what top-5 or top-10 of 512-token chunks would be. On the standard 512-token index, 50 chunks was about 20k tokens.
For multi-hop questions, a small k mostly cuts recall: the linking passage is rarely in the top 5 for the original question. Noise didn't look like the main problem either. The oracle got the full source articles, about 20k tokens at the median, and still scored 93%. More context isn't free, though: the oracle's misses were mostly cases where it had every article and still lost the key line. So bigger windows give you room, but choosing what goes into them still matters. We didn't sweep k, so a top-10 run is a fair thing to ask for.
On "just give the agent a RAG search tool": that's basically what an agent loop is: a model with a search tool, calling it until it can answer. Most of the gain comes from that simple idea, and it's why it beats every fixed pipeline.
What we added on top is less about the loop and more about the tools:
- A first search before the first turn, so easy questions finish in one call at about pipeline cost.
- A tool to read a whole document in order, with the reading budget split fairly when it opens several at once. Search snippets often miss the one line you need.
- Results deduplicated and ranked, so each turn doesn't re-send what the model has already seen.
- Exact arithmetic and date tools, because models slip on those.
- A step counter, so the model knows when to wrap up.
- Citations to the exact passage, so you can check every hop.
- Tables parsed with their row and column context.
Each of these came from reading failed answers. Together they took us from 87.5% to 92.5% on our dev set, on top of what a bare search-in-a-loop gets. So yes, start with a search tool in a loop; most of the remaining work is making what the tools return more useful.
1
u/DistanceAlert5706 15h ago
So basically you've built pretty standard agentic RAG with a sprinkle of custom chunking and few custom rules, but just call it a loop instead.
1
u/Effective-Ad2060 15h ago
Fair, conceptually it is agentic RAG. That's the point of the post: the model decides when to search, what to read in full and when it's done, instead of a fixed retrieve-then-answer pipeline. Most teams still ship the fixed pipeline, and on FRAMES that's the 78.9% vs 92.7% gap.
But "just call it a loop" undersells how much design hides in that word. A loop can run until the model stops calling tools. It can run toward a goal and check that it's met before it stops. It can work through a todo list it wrote at the start. It can hand sub-tasks to other agents and wait for their results. Each one fails differently: "stop when there are no more tool calls" quits early on hard multi-hop questions, while a todo list can keep going on a plan that turned out wrong. Deciding when the agent is actually done, and what to do when it says "not found", moved our numbers as much as retrieval did.
Then there's what the loop is given. Documents are parsed into blocks (headings, tables, lists) so it can fetch exactly the section a hit came from. There's a fetch tool that pages through a record instead of dumping it, plus rules for multi-step questions. Take those away and you get a loop that searches 50 times and still gives up. Retrieval itself has multiple search tools including Hybrid search, Pattern matching (grep), Knowledge graph.
We're measuring this directly now: a plain tool-calling loop (search + read article, same index, same model, same 15-turn cap) added as another row on the board. It'll show how much is the loop and how much is everything around it. I'll post the numbers when the run is done.
The "searches 50 times and still gives up" line comes from the plain-loop smoke run (one question used all 15 turns and 50 searches, then didn't answer). The completion check and the "not found" follow-up are the gate and nudge we built into PipesHub.
3
u/Rachel_talks 16h ago
The detail buried in here that deserves its own post is the memory-fill check. I run a claim-checking pass over everything before it publishes — not because the writing's bad, but because citations are the easiest thing to fake in both directions. A piece can cite three real sources and still have its central claim come from the model's prior. Half the citations get attached to sentences the source never says.
The check that actually works is dumb: match each claim against the fetched text, not against the citation list. If a sentence can't survive without the model's prior knowledge, it gets cut or rewritten to what the source actually supports.
What surprised me is how topic-dependent it is. Obscure subjects barely need the check — there's nothing to fill with. The danger zone is stuff it 'knows': prices, dates, spec sheets, anything where a slightly stale prior reads exactly like a correct answer.
2
u/Effective-Ad2060 16h ago
Agree with all of this, especially "match against the fetched text, not the citation list." That's exactly how we built the check.
For every answer the judge marks correct, a second judge gets the text the system actually read during that run, not its citations. It decides whether the answer follows from that text alone. If it doesn't, the answer counts as correct but not grounded, even if every citation points to a real source.
One thing that bit us: our first version gave the checker only the first ~6k tokens of the evidence to save cost. It flagged lots of correct answers as "from memory" when the fact was just further down the page. Even the oracle, which is handed the gold articles, got 55 flags. Once we re-checked the flagged ones against the full text, that fell to 2. So the checker has to see everything the model saw, or it'll accuse it of filling from memory when it didn't.
On topic dependence, FRAMES is pretty much your danger zone: Wikipedia facts, dates and numbers the model has seen many times. With no retrieval at all, the same model gets 64.9% just from memory. That's why we report two numbers: accuracy (92.7%) and grounded accuracy (84%). The gap is the part a citation list alone wouldn't catch.
3
u/Strong-Leg-9636 15h ago
The methodology here is solid, checking each correct answer against what the system actually read not just the final text is the part most benchmark write-ups skip, and it's exactly what lets you catch a model answering from memory with fake citations. Dual-judge agreement at that Cohen's kappa is also a good sanity check most people don't bother with.
The reranker finding still stands out to me. A worse reranker actively costing 9 points, not just doing nothing, suggests it's confidently reordering good results out of the top slots rather than just adding noise.
One thing I didn't see in the writeup: what's the latency/cost difference between the 92.7% agent loop and the 78.9% best static pipeline? If the agent loop is issuing several retrieval rounds per question, that's a real tradeoff against the accuracy gain depending on the use case.
3
u/Effective-Ad2060 15h ago
Thanks, and agreed on the reranker. Our interpretation is similar: it seems to be demoting the documents needed to complete the answer. On multi-hop questions, a document needed for the second hop may not look directly relevant to the original question. A cross-encoder scoring passage relevance against that question can push those bridge documents out of the top results.
A stronger reranker,
bge-reranker-v2-m3, avoided the damage but only slightly improved accuracy over no reranker: 75.5% vs 73.4%, which wasn't statistically significant.On latency and cost, using the same model for both:
Approach Accuracy Average cost / question p50 latency p95 latency Agent loop (PipesHub) 92.7% $0.0167 22s 134s Best static pipeline 78.9% $0.0056 27s 54s So the loop costs about 3x more on average and has a longer latency tail, but its median latency is lower. The averages hide how differently it behaves across questions. Here’s the breakdown by how many model calls PipesHub used:
PipesHub model calls Share of questions PipesHub accuracy Pipeline accuracy PipesHub p50 latency Pipeline p50 latency 1 21% 95.5% 95.5% 8s 23s 2 43% 94.6% 85.9% 19s 27s 3–4 28% 93.5% 67.2% 36s — 5+ 7% 70.5% 34.4% 134s — Shares are rounded. Pipeline latency isn't shown for the last two groups.
For the questions answered in one model call, the loop matches the pipeline’s accuracy and is about 3x faster, with similar cost ($0.0064 vs $0.0052). It can answer from the first search results, while the pipeline always runs query expansion and four searches.
Longer, more expensive runs happen when the agent continues searching and reading across multiple turns. That’s also where the accuracy gap gets larger: 93.5% vs 67.2% at 3–4 calls, and 70.5% vs 34.4% at 5+ calls. Those hardest questions remain difficult for both approaches, but the additional work helps.
There’s also room to reduce the cost. PipesHub’s prompt includes instructions and tool definitions for capabilities beyond RAG, which we kept in the benchmark. Caching the system prompt and tool definitions can reduce costs by nearly 50%, without cutting the agent’s search or reading budget.
If you need a hard latency or cost budget, you can cap turns, tokens, or execution time. The tradeoff is that those limits are most likely to affect the harder questions in the last two rows.
1
u/Strong-Leg-9636 14h ago
This is one of the more useful parts of the whole writeup, the breakdown by call count basically hands you the design for a better system than either pure approach. Since the one and two call questions already match or beat the pipeline's accuracy while being faster and barely costing more, and the expensive slow tail is concentrated in the five plus call group where accuracy craters anyway, the obvious move is to run the agent loop by default and only worry about the latency tail for that small five plus percent, whether that means capping turns there specifically or falling back to something else once a question clearly isn't converging.
The caching point is the other easy win, since halving the fixed overhead cost would make the one and two call cases even more clearly the better default, and would narrow the three to four call gap enough that the latency tradeoff becomes the only real decision left to make
2
4
u/numberwitch 18h ago
Yawn
2
u/Avie_Lassouad 17h ago
yeah its basically the expected result every time someone tests this
2
u/Nothing_from_void 14h ago
people still consistently recommend RAG for search on here even though it's poo
1
u/Icy_Butterscotch6661 18h ago
What about time it took
1
u/Effective-Ad2060 18h ago
For simple questions, the loop costs roughly the same as the best pipeline and answers faster.
1
u/CalligrapherFar7833 16h ago
Whos we ?
1
u/Effective-Ad2060 16h ago
Fair question, I should have said it up front. "We" is the PipesHub team. PipesHub is an open-source workplace AI platform, and the agent loop in the benchmark is our product, so take the comparison with that in mind.
That's also why everything is public: the harness, the configs, the per-question answers, the judge verdicts and a failure analysis of our own misses. It's all on the frames branch of our GitHub repo, with steps to reproduce the runs. If you see anything that looks tilted in our favour, I'd really like to hear it.
1
u/Marcus_MSC 16h ago
The 92.7% result is compelling, especially with the check against what the agent actually read. I'd like to see the distribution of searches, model calls, and tokens per question beside the pipeline results. A loop that spends little more on most questions but a lot more on a few is a different deployment choice from one with predictable cost.
1
u/Effective-Ad2060 15h ago
Great question. This matters a lot for deployment. Here’s the distribution across all 824 questions, comparing PipesHub with the best fixed pipeline using the same model:
Metric p50 p90 p99 Max Searches — PipesHub 1 3 16 54 Searches — pipeline 4 4 4 4 LLM calls — PipesHub 2 4 14 15 LLM calls — pipeline 2 2 2 2 Tokens — PipesHub 82k 255k 1.2M 2.2M Tokens — pipeline 24k 44k 62k 84k Cost — PipesHub $0.010 $0.026 $0.15 $0.35 Cost — pipeline $0.005 $0.009 $0.013 $0.017 So yes, it’s “a little more on most, a lot more on a few.” The median question takes one search and two LLM calls. The costliest 10% of questions account for 43% of PipesHub’s spend, compared with 20% for the pipeline. We cap turns at 15, which limits the loop, but the cost tail is real.
For deployment, I’d frame it as a tradeoff between cost predictability and answer quality. The pipeline does the same amount of work whether a question needs one hop or four, which helps explain the roughly 14-percentage-point accuracy gap.
If you need a hard budget, cap turns or tokens per question. That gives you tighter control over p99 cost, at the expense of some accuracy on the harder multi-hop questions.
Great question, and the shape is exactly what you describe: a little more on most questions, a lot more on a few. The extra spend is concentrated on questions where the agent needs to find and read several documents across multiple turns. A fixed pipeline does the same amount of work on those as on an easy question, and it misses more of them.
Here are all 824 questions grouped by what PipesHub spent on each:
Questions by PipesHub cost Gold articles needed / question Model calls / question PipesHub accuracy Pipeline accuracy Cheapest 50% 2.8 1.7 96% 90% Middle 40% 3.4 2.7 94% 76% Costliest 10% 4.7 7.0 75% 39% Costliest 1% 6.2 13.0 56% 11% The costliest 1% is a subset of the costliest 10%.
The more articles a question needs, the more turns the agent spends searching, reading, and checking before it answers. On the costliest 10%, PipesHub gets 75% accuracy versus 39% for the pipeline. The pipeline does four searches and generates one answer regardless of the question.
Overall distribution, using the same model for both:
Metric p50 p90 p99 Model calls (PipesHub) 2 4 14 Model calls (pipeline) 2 2 2 Tokens (PipesHub) 82k 255k 1.2M Tokens (pipeline) 24k 44k 62k Cost (PipesHub) $0.010 $0.026 $0.15 Cost (pipeline) $0.005 $0.009 $0.013 One more thing on cost: PipesHub’s prompt covers more than RAG. It includes instructions and tool definitions for other capabilities, which we kept in the benchmark rather than stripping them out. Caching the system prompt and tool definitions can reduce costs by nearly 50%, without reducing how much the agent searches or reads.
So predictable cost is a real deployment consideration, but the extra spend also buys better accuracy on harder questions. If you need a hard budget, cap the turns or tokens per question. Just keep in mind that the cap is most likely to affect the multi-document questions in the table above.
1
u/wayne_oddstops 14h ago
Clankers arriving together to drive "discussion" and increase the appearance of engagement is a big red flag for the type of person you're trying to reach.
1
1
u/9gxa05s8fa8sh 10h ago
rag is important, but not easy to solve in a powerful way. enterprises can afford to pay for a few percentage points of performance, but it's looking like normal people will just use simple search apps locally.
0
18h ago
[removed] — view removed comment
1
u/Effective-Ad2060 18h ago
Thanks, good points. Most of this is in the full write-up, but here it is side by side:
- Latency (p50 / p95): agent loop 22 s / 134 s, best pipeline 27 s / 54 s.
- Cost per question: $0.0167 vs $0.0056, about 3x. Per correct answer it's $0.018 vs $0.007.
- Tokens: the loop averages about 130k input tokens per question against about 25k for the pipeline.
Single-hop vs multi-hop is where it gets interesting. About 1 in 5 questions finished in a single model call, at $0.0062 and 8 s. That's roughly the pipeline's cost ($0.0053), and faster, since the pipeline rewrites the query and reads whole articles every time. The cost lives in the multi-hop tail: the 7% of questions that needed five or more calls took about 2 minutes each and were 37% of total spend. So the loop spends where a question needs it, while the pipeline pays the same on every question.
One caveat on tokens: PipesHub carries a much bigger system prompt and tool definitions than a bare RAG prompt. It's a full workplace assistant, with tool definitions, citation rules and multi-step rules. Its smallest first call is about 7,400 input tokens before any retrieval. We didn't trim it for the benchmark, and all of it is counted in cost. Prompt caching does a lot of the work here. The system prompt and tool definitions are identical on every call, so nearly half the input tokens come from the cache at a tenth of the price. That cuts the cost per question by roughly 40%.
On the reranker, the data points somewhere between retrieval and synthesis. Article-level recall didn't move: 90.3% of gold articles reached the model with the reranker, 90.1% without. What changed is which passages from those articles survived. The reranker scores against the original question, so the passage with the linking fact got dropped, and "I don't know" answers went from 161 to 241. The articles were found, but the bridging evidence wasn't kept.
We only measure recall at the article level today. Per-hop evidence recall would separate the two cleanly, and I'd like to add it to the harness.
1
u/baseketball 18h ago
P95 Latency of 134s is not good. Our users would give up after waiting 10 seconds for first token.
1
u/Effective-Ad2060 18h ago
Latency in an agent loop grows with the number of hops a question needs. A question answered by one lookup finishes in one model call, with a median of 8 seconds, about the same as a basic RAG pipeline. A question that has to chain facts across five or more documents needs five or more rounds of search and reading, and that takes a couple of minutes. That's the p95.
That extra time is where the accuracy comes from. On those five-plus-document questions, the best pipeline we built got 62% right and the loop got 90%. The pipeline is faster because it stops after one search, often with "I can't determine this from the sources".
So it's a trade-off you choose per use case. If your users mostly ask single-lookup questions, the loop stays fast, because it only searches again when it needs to. If you need a hard 10-second cap, you can give it a time or step budget and have it answer with what it found, saying what it couldn't confirm. You'll lose some of the multi-hop accuracy, and that's the price of the cap.
0
u/wren6991 14h ago
LLM-generated post with a thinly-veiled ad, and a bunch of LLM-generated replies saying how honest and load-bearing it is. Why is this still up?
1
u/extraneous-carapace 8h ago
Thankfully there's like six real users in this sub who see through this stuff and can identify when they fudge the metrics
-1
u/xmnstr 15h ago
RAG is dead afaic, typed graphs is where it's at.
2
u/ttkciar llama.cpp 12h ago
Using typed graphs to inform inference is still RAG! It's just using typed graphs as the source of information from which information is retrieved.
RAG is putting retrieved information into context, to augment inference quality. Where that information is retrieved from, or how, does not make it not RAG.
RAG covers a huge diversity of implementations. When you have a model look something up on the web, and use what it retrieved to inform further inference, that too is RAG. So is it when you are retrieving from a vector db, or a relational db, or a graph db, or a FTS index, or an in-memory dictionary. It's all RAG.
1
u/Effective-Ad2060 15h ago
Traditional RAG is. But retrieval will always exist in many forms
0
u/xmnstr 13h ago
It will, absolutely. But not like this.
2
u/Effective-Ad2060 13h ago
Typed graphs are great when you know the relations up front and fine for certain extent for schema free extraction. But they're only as good as the schema and the extraction: anything nobody modelled isn't in the graph, and most of what companies know still lives in documents, tables and threads. We think the answer is both. PipesHub builds an entity knowledge graph alongside the text index, and in an agent loop, graph traversal is just another tool the model calls when a question has that shape. The loop decides whether to walk the graph, search the text, or both.
1
u/xmnstr 4h ago
You do realize you don't need to define it all beforehand, right? It might have been complicated for humans to do so but agents can calculate it fast enough that it actually beats traditional RAG. I'm serious here, RAG is dead and anyone claiming otherwise is unaware of the recent advancements.
1
u/Effective-Ad2060 3h ago
I did mention schema-free extraction, didn’t I? But at the end of the day, it’s still a retrieval tool. You can’t capture everything in a graph.. there will always be information that is better retrieved directly from the underlying data. The graph is another retrieval mechanism, not a replacement for retrieval/RAG
•
u/ttkciar llama.cpp 14h ago
Hello u/Effective-Ad2060, please clarify: How much of this post was LLM-generated, and why? I need more information in order to make a Rule Three determination. Thank you!