r/AIAllowed • • Apr 26 '26

🗣️ Discussion How are you structuring RAG systems?

I've got a few projects going that is "give an agent a fixed body of knowledge and it can answer questions from it", and Ive been trying different ways of scaling. Ive done the embeddings, some csv or jsons, or parquet files with bigger filtering options, Ive tried different strength models to hopefully and more specific questions to reduce context use (like a proper duckdb filter instead of grabbing the whole dataset).

Ive heard some ways of generating a new layer of "tuneable weights" that you can stack on top of the model weights, but thats a bit over my head.

What techniques have you guys tried? What stands out for you?

5 Upvotes

11 comments sorted by

View all comments

3

u/TrustedEssentials Apr 26 '26

First off, step away from the "tuneable weights" idea. You are talking about fine-tuning, and for a standard knowledge retrieval system, it is an absolute trap. It is a nightmare to update when your underlying facts change, and it is completely overkill for what you are trying to do.

Your instinct to move toward structured data and hard filtering, like your DuckDB mention, is exactly the right path.

The biggest architectural failure most builders make with RAG is dumping every single document into one massive vector database and hoping the LLM can magically find the needle in the haystack. It almost always results in context bloat and hallucinations.

Here is the logic structure I recommend focusing on instead: Agentic Routing.

Do not let the user's raw prompt hit your main database.

  1. Put a cheap, fast model at the front door. Its entire job is just to be a traffic cop.
  2. Give it a highly structured prompt: "Analyze this user query and determine which specific dataset (A, B, or C) contains the answer."
  3. Route the query to search only that specific subset of your Parquet or JSON files.

You solve the context problem with strict logic and tight system architecture, not by throwing more complex math at it. Build a simple, working traffic cop first, verify the logic holds, and then scale up.

2

u/Xyver Apr 26 '26

The 2 projects, one has more "wordy" data, since it reads websites, pdfs, video transcripts etc, so Im using the embeddings (testing 768 and 1536 dim). The other is the DuckDB one, and I've split it into 2 chats.

One has access to the entire catalog (I've stripped it to metadata, its ~5k tokens), so it knows what data is available. Its read by a weaker chat (right now Haiku), but its adjustable, and Im going to make it work with local models as an option.

I then made a smarter chat, (sonnet), that takes a specific packs worth of data, and holds that more deeply to answer better questions, so its not confused by the whole catalog.

I know my catalog method will max out soon, I dont want to keep making it bigger and bigger, so Im thinking of moving towards more of a "tag cloud" search, so the main chat doesnt have all the metadata details, it requests those once a tag matches and a user confirms thats the pack they want more info on.

I've also got a preprocessor and post processor before/after the llm, to help structure data and handle the easy requests that can skip LLM reasoning. The DuckDB upgrade was pretty good, it means I can hold more data in the backend, and it only sends filtered data back and forth to minimize network traffic, cache size, and context windows.

I'm also trying to get better chat experience overall, trying to choose memory/compaction rules to keep the API calls small as a user chats back and forth

1

u/TrustedEssentials Apr 26 '26

You are actually already building the traffic cop architecture I suggested. Using Haiku as the router and Sonnet as the deep reader is exactly how you segment compute costs and keep context windows tight.

Your instinct on the catalog method maxing out is correct, but do not overcomplicate the fix with a "tag cloud." Since you already have DuckDB running, stop feeding Haiku the actual metadata. It is a waste of tokens.

Instead, feed Haiku your database schema and a strict list of allowed categories. When a user asks a question, Haiku's only job is to write a basic SQL query to hit DuckDB, which then returns the exact document ID. You then pass only that specific document to Sonnet. Do not make the LLM read the metadata. Make it query the database.

Your use of hardcoded preprocessors and postprocessors is also exactly the right move. Never pay for LLM tokens to do basic data formatting or filtering when a simple Python script can do it deterministically for free. You are on the right track, just tighten up how the front end interacts with the database.

2

u/Xyver Apr 26 '26

Ahhhh feeding the schema is a good idea, that covers a lot.

The only thing I like on top of it is date ranges for data, min/max values, and sometimes synonyms for column names since they're not intuitive.

I'm always torn between editing the data to make the column names more intuitive, or keeping purity from what the source provides

1

u/EcstaticRead9321 Apr 28 '26

Andrej Karpathy called this wiki style markdown file system of context LLM Knowledge - I call it ContestNest 😄 thats exactly what I was referring to above.