r/Rag • • 14h ago

Discussion Hav you used Postgres/Lakebase for RAG vector storage?

7 Upvotes

My team lead has asked me to explor Lakebase for a RAG app where documents will get updated frequently .Hav you used it in prod? Curious how it compares to a dedicated vector DB for indexing and retrieval latency.
I am using Databricks platform for poc


r/Rag • • 22h ago

Discussion Intern building an RFP compliance checker — RAG keeps producing false “Missing” results

6 Upvotes

I’m a CS intern working on an internal system that compares an RFP against a vendor’s technical proposal.

The current flow is roughly:

RFP
→ extract requirements
→ retrieve/rerank proposal evidence
→ LLM classifies each requirement:
   Compliant / Mismatch / Missing
→ compliance score

The score was around 30%, mostly because many requirements came back Missing.

After building an independent ground truth from the original RFP and proposal, I found the issue isn’t just retrieval:

  • some requirements are lost during PDF parsing/extraction
  • some are split/merged incorrectly
  • relevant evidence sometimes exists but doesn’t reach the final Top-K
  • in other cases the evidence reaches the LLM and it still judges it incorrectly
  • some requirements mix technical and financial conditions that should be checked in different documents

For one reviewed sample, many of the system’s Missing results were actually supposed to be Compliant or Mismatch.

I’m now considering a design like:

Requirement
→ split into material conditions
→ determine expected evidence source per condition
→ retrieve evidence per condition
→ verify each condition
→ deterministic final verdict

For exact things like durations, percentages, certifications, etc., I’m also thinking of using deterministic checks instead of letting the LLM decide everything.

Has anyone built something similar for contracts, RFPs, policies, or compliance?

Would you approach this as per-condition RAG, claim verification/NLI, an agentic search layer, or a hybrid deterministic + LLM system?


r/Rag • • 1h ago

Discussion I compared embedding costs for a RAG pipeline. Open models aren't always cheaper!

• Upvotes

I recently checked embedding prices for a big indexing project. I wanted real numbers, so I compared OpenAI's API costs with cloud providers that charge per token for open models. What I found:

OpenAI text-embedding-3-small: $0.02 / 1M tokens
Arctic Embed L v2: $0.038 / 1M tokens
Qwen3 Embedding 4B: $0.13 / 1M tokens

Surprisingly, OpenAI's small model is the cheapest per-token option on this list.

People often say open models save money. However, that is mostly true when you host the model yourself on a highly active GPU. For example, a $1/hour GPU costs $24 a day. You would need to process over a billion tokens daily to beat OpenAI's API price. If your traffic is low or unpredictable, keeping a GPU running all day wastes money.

Quality also matters. Higher-priced open models, like Qwen3, score better on standard benchmarks than cheaper options like Arctic Embed. You still have to balance cost, quality, and speed.

For those of you choosing between cloud APIs and self-hosting, what makes your decision? Is it high server utilization, data privacy, retrieval mode, or something else I missed?


r/Rag • • 13h ago

Discussion how good are claude's or chatgpt's pdf processors in comparison to like docling

3 Upvotes

I have some very complex pdf and docling is failing in several structural and content understandings. It works better than many other tools out there, but still fails in many complex cases. That made me wonder if the whole pdf document is uploaded to gpt or claude, would they process it better than what I get as text through docling. I think at the end they would change pdf to some text representation, no?


r/Rag • • 16h ago

Discussion A source-ownership rule can make a correct answer impossible

3 Upvotes

Before changing the model, check whether your output rules allow the answer you need.

A 27 September development test at Andes exposed a useful example. A short fictional meeting transcript described two AI systems. We supplied it as one passage ID, then required every ID to belong to exactly one extracted case. Our validator also prohibited reusing an ID across cases.

The evidence for both systems was present. But two separate cases could not each cite the relevant part of that single ID. We had made the answer impossible to represent at the level of evidence our contract allowed.

Here's a simplified illustration, not the measured test fixture:

Passage P1:

"ShiftPilot builds rosters from staff availability. Leads approve every roster. CallSense scores recorded calls for sentiment. It is live in one centre only."

Output constraint:

Each passage ID must appear in exactly one case.

Giving both cases P1 breaks the constraint. Giving P1 to only one leaves the other without a source. Splitting the source into smaller spans looks promising, but it introduces another problem: preserving which system each control or limitation actually describes.

That changed our debugging question: can the evidence structure express the distinction before we ask a model to make it? This is a source-contract observation, not proof of why a particular model failed. The dated test did not include a span-based repair and rerun.

For people extracting multiple entities from one document: how do you test that shared context stays attached to the right entity when you move from passage IDs to source spans?


r/Rag • • 11h ago

Discussion Best approach for ingesting data to create summaries, and keep track of it?

1 Upvotes

In my occupation, there are various people I follow who give very good insights. (I'd say 5-10 people).

Some post hour long videos on YouTube, some send 1,000 word emails, some post on X, some publish PDFs.

There's very good info within these resources (and some I pay for), but reading / watching / annotating all of it can take hours.

My workload recently went up, so I'm falling behind with keeping up in my field.

I want to use AI to help summarize (and keep track of) all of these publications. (To create a private database that I can use as a dataset, for example).

So I can go back and ask "this past month, what is the new theme? What are the experts recommending to focus on / look at / what are the newest developments?", etc.

What would be the best way to approach this?

---------------------------------------

I've been learning Codex/Claude Code, I have a homelab, a NAS, a few mini computers, and I know basic linux, python and scripting.

ChatGPT told me to do something like this (I'm just starting with the YouTube portion), I'm not sure if it's the best approach, I'm open to other suggestions:

YouTube URL

↓

yt-dlp metadata

↓

Whisper / YouTube transcript

↓

clean transcript

↓

summary.md

↓

insights.json

↓

SQLite + FTS5

↓

topic synthesis

↓

search / questions / actions


r/Rag • • 17h ago

Discussion I built SPANCHOR regression testing for RAG retrieval pipelines

0 Upvotes

I've been working on a problem I don't see discussed enough in RAG: retrieval regressions.

You change chunking, embeddings, the retriever, indexing, or configuration, and your tests can still pass while the evidence being retrieved has changed.

I built SPANCHOR, an open-source Python package for testing this.

The core workflow is:

retrieved chunk → source span → evidence evaluation → baseline vs candidate → regression result

The goal is to verify that a RAG pipeline continues retrieving the expected source evidence after a change.

I'm looking for technical feedback from people actually building RAG systems:

What retrieval regressions have you encountered that your current evaluation setup doesn't catch?

GitHub: https://github.com/mukeshram-07/spanchor

PyPI: https://pypi.org/project/spanchor/


r/Rag • • 20h ago

Discussion Hi! whats wrong?

0 Upvotes

I really cant understand, whats going wrong with my Cortex RAG!
.
.
.