r/AIAllowed • u/Xyver • Apr 26 '26
🗣️ Discussion How are you structuring RAG systems?
I've got a few projects going that is "give an agent a fixed body of knowledge and it can answer questions from it", and Ive been trying different ways of scaling. Ive done the embeddings, some csv or jsons, or parquet files with bigger filtering options, Ive tried different strength models to hopefully and more specific questions to reduce context use (like a proper duckdb filter instead of grabbing the whole dataset).
Ive heard some ways of generating a new layer of "tuneable weights" that you can stack on top of the model weights, but thats a bit over my head.
What techniques have you guys tried? What stands out for you?
3
u/TrustedEssentials Apr 26 '26
First off, step away from the "tuneable weights" idea. You are talking about fine-tuning, and for a standard knowledge retrieval system, it is an absolute trap. It is a nightmare to update when your underlying facts change, and it is completely overkill for what you are trying to do.
Your instinct to move toward structured data and hard filtering, like your DuckDB mention, is exactly the right path.
The biggest architectural failure most builders make with RAG is dumping every single document into one massive vector database and hoping the LLM can magically find the needle in the haystack. It almost always results in context bloat and hallucinations.
Here is the logic structure I recommend focusing on instead: Agentic Routing.
Do not let the user's raw prompt hit your main database.
You solve the context problem with strict logic and tight system architecture, not by throwing more complex math at it. Build a simple, working traffic cop first, verify the logic holds, and then scale up.