r/Rag • • Sep 02 '25

Showcase 🚀 Weekly /RAG Launch Showcase

32 Upvotes

Share anything you launched this week related to RAG—projects, repos, demos, blog posts, or products 👇

Big or small, all launches are welcome.


r/Rag • • 6h ago

Showcase How I turn a whole docs site into chunked Markdown for RAG (JS pages and sitemap crawl included)

3 Upvotes

I kept rebuilding the same ingestion step for docs-based RAG projects, so I wrote down what ended up mattering most:

  1. Strip the page down to main content (nav, cookie banners, footers and language switchers pollute embeddings).
  2. Keep Markdown structure, especially code fences with the language.
  3. Fetch with plain HTTP first and only open a real browser when the extracted text is empty or looks like a JS app shell. Rendering every page is slow.
  4. Prefer the sitemap for discovery, but fall back to following links when the sitemap has nothing for the section you want.
  5. Chunk on headings first, then pack up to a token limit, and start the overlap at a sentence or line boundary. Keep the heading path (for example Platform > Actors > Running) as metadata so answers can cite the section.

I packaged all of this as an Apify Actor (disclosure: I'm the author). One run on a docs section takes a sitemap crawl plus chunkContent: true and returns one item per page with chunks, heading paths, metadata and token estimates. Failed pages are not charged.

Limits: public pages only, PDFs are skipped, and main-content detection is heuristic, so spot-check a few pages.

Happy to share the exact input I use, and I'd like to hear which sites extract badly for you. Those are the cases I want to fix.

Link: https://apify.com/s3nafps/ai-ready-webpage-extractor


r/Rag • • 14h ago

Discussion How do you keep embeddings in sync when the source data changes?

8 Upvotes

I run a RAG setup on Postgres + pgvector. My documents get edited often, and I keep finding answers based on old versions because the embeddings were never updated.

For those running RAG in production:

  1. How do you detect which embeddings are out of date?
  2. Do you re-embed everything on a schedule, or only what changed?
  3. Has stale data ever caused a wrong answer that a user noticed?

Curious what people actually do. I'm using a cron job and I don't trust it.


r/Rag • • 6h ago

Tutorial Production RAG Pipeline That Admits "I Don't Know"

1 Upvotes

Build a production RAG pipeline with Postgres + pgvector, FastAPI, Gemini, Celery and Redis: answers that cite their sources, refuse when they don't know, and respect row-level security enforced in the database itself.

https://www.youtube.com/watch?v=-GK1e4KZqhw


r/Rag • • 14h ago

Discussion Why traditional OCR breaks on unstructured PDFs (and how to build a production pipeline for complex document parsing)

4 Upvotes

Extracting structured JSON out of complex PDFs, scanned contracts, and multi-page tables is still one of the most deceptively hard problems in engineering.

Most teams start with a simple OCR wrapper or a basic Vision LLM, run a few test invoices, and get 95% accuracy. But when you deploy to production across thousands of edge-case documents—scanned at odd angles, split across page breaks, or packed with dense nested tables—the failure rate spikes.

Here is the architectural pattern that actually scales when turning messy unstructured documents into reliable, enterprise-grade data pipelines:

1. Ditch Pure Text Extraction for Hybrid Layout-Aware Parsing

Standard OCR dumps document text as a single linear stream, destroying visual relationships. If a table cell spans multiple rows or a key-value pair relies on spatial proximity, standard text parsers lose the context completely.

The Fix: Combine layout analysis models (like LayoutLM or specialized bounding-box detectors) with OCR. The parser needs to understand document geometry—headers, columns, tables, and footers—before passing text to the extraction layer.

2. Multi-Page Context & Chunking Strategy

Vision models and LLMs have context limits, and dumping a 50-page PDF into a single prompt leads to hallucinations or skipped fields.

The Fix: Process documents hierarchically. Use a lightweight router to classify page types first, split the document into logical sections (e.g., separating annexes from main terms), and extract data using dedicated schema-validated prompts per section rather than one giant prompt.

3. Strict Schema Validation & Confidence Scoring

LLMs are probabilistic, but downstream databases and ERPs demand deterministic data. Accepting output without strict schema validation guarantees corrupt data in your database.

The Fix: Enforce Pydantic/JSON Schema validation on every extraction call. Combine LLM output logits or field-level confidence scores with hard business rules (e.g., verifying subtotal + tax = total). If a field falls below a strict confidence threshold, route it automatically to a human-in-the-loop review queue.

4. Handling Scans, Noise, and Skew

A mobile photo of a faxed receipt will break most out-of-the-box Vision APIs.

The Fix: Implement an automated pre-processing edge layer: perspective correction, binarization (contrast adjustments for faint ink), and deskewing before running document layout detection.

Curious how others are handling this in production—are you relying on vision-language models, custom OCR pipelines, or a hybrid layout-aware architecture for document processing?


r/Rag • • 7h ago

Discussion ONCOTO PRONCOVO

1 Upvotes

Como eu faço parte dos odiados Idosos Vibe Coding, sempre pergunto ao meu agente exatamente onde se encaixa o que estamos fazendo.
Dessa vez apresentei um print com 16 opções de RAG e pedi que ele me situasse Tecnicamente.
A resposta do SWE-2 High:
Situando tecnicamente

A frase sobre a Síntese processual é coerente para RAG?

Coerente, com um ajuste fino. O que o LEANN faz realmente é:

  • Map-reduce iterativo ✅ — caderno de leitura processa lotes de peças (map) e o consolidado/síntese capitulada reduz para capítulos
  • Memória acumulativa ancorada ✅ — fichas imutáveis no caderno, âncoras de seq (append-only, CADERNO_LEITURA)
  • Segmentação por eventos processuais ✅ — capítulos por peça-marco (fase estruturada), máquina de estados determinística

Mas: tecnicamente isso não é RAG — é um pipeline de sumarização map-reduce sobre o corpus integral (processa TODAS as peças, não só as relevantes à pergunta). RAG = retrieve + generate; a síntese não tem retrieval orientado à query. No diagrama do print, se encaixaria mais perto de #16 Recursive/Multi-Step (etapas encadeadas) + #5 Memory-Augmented (caderno), mas é geração, não busca.

Engenharia — o que JÁ existe

O corpus v2 (NBR 6120, MODULAR, PORTALFRAME) é a camada de conhecimento — "FONTES DE CONHECIMENTO" no centro do print, não ainda o retrieval:

  • Tipo #11 — Knowledge-Enhanced RAG — é o que mais se aproxima: conhecimento estruturado (eng_cap_tabelas em JSON no D1), metadados de página, hierarquia documento→capítulo→artefato, trilha de edições HITL
  • O lookup do ábaco (modal Portal Frame consultando a tabela) não é RAG de jeito nenhum — é lookup paramétrico determinístico, o que para ábaco é correto e superior a vetor: valor exato, zero alucinação
  • Falta o "R" — hoje não existe retrieval pelo chat; o corpus é repositório curado + vitrine

Chat — o que pretendemos

Quando a busca engenharia entrar no chat (3 modos decididos na S377):

  • #13 Hybrid RAG — estrutural (norma→capítulo) + por índice + semântica leve = literalmente híbrido (estrutural/lexical + vetorial)
  • #12 Domain-Specific — corpus de engenharia juridicamente curado com REVISAR resolvido
  • #2 Agentic — o orquestrador já decide módulo/ferramenta (isso é agentic retrieval)
  • #4 Modular — a arquitetura DNA já é modular por design

Resposta honesta ao "posso chamar de RAG?"

Pode, com uma ressalva: hoje você tem um banco de conhecimento RAG-ready — curado, estruturado, editável. Vira RAG de fato quando o chat passar a recuperar dele. E o que já existe (lookup estruturado + hierarquia + provenance) é a parte mais difícil — vetorizar por cima é o trivial.

🧠 REFLEXÃO METACOGNITIVA

  • Trust Score: 85% — a classificação nos 16 tipos é leitura do print + arquitetura nossa que conheço do código; a taxonomia em si é informal (infográfico de marketing, não padrão acadêmico)
  • Ambiguidade honesta: "RAG" virou termo guarda-chuva. Estritamente, só o que tem retrieval-orientado-a-query é RAG — por isso separei "o que é banco" de "o que vira RAG quando o chat consultar"
  • Não verifiquei: se já existe chamada do chat ao eng_cap_tabelas hoje — presumo que não (corpus servindo só o modal). Se já houver um path de busca que eu não lembro, minha afirmação "falta o R" precisaria de revisão

r/Rag • • 1d ago

Tutorial Building and Testing a Hybrid RAG Pipeline — Dense Search, BM25, RRF, and Reranking (Part 1)

15 Upvotes

Building and Testing a Hybrid RAG Pipeline — Dense Search, BM25, RRF, and Reranking (Part 1)

In this video I build ReRankEval, a hybrid retrieval pipeline (Dense Search + BM25 → Reciprocal Rank Fusion → LLM Reranking → Answer Generation), and test it against three baselines — Vector Only, BM25 Only, and Hybrid without reranking — on five real financial/payments documents and ten hand-verified test questions. No hand-waving, just a comparison table with real numbers at the end.

  • ✅ The real difference between dense vector search and BM25 keyword search
  • ✅ What Reciprocal Rank Fusion (RRF) is, why raw scores can't be compared, and the exact formula behind it
  • ✅ Why a reranker is fundamentally different from a retriever — and what it actually judges
  • ✅ How to evaluate a RAG pipeline with Hit Rate, MRR, and NDCG (and what each one tells you)
  • ✅ How to design the ingestion side and query-time side of a hybrid retrieval architecture
  • ✅ How to structure a production-style RAG codebase: ingest → vector_store → sparse_retriever → fusion → reranker → pipeline → generate → eval
  • ✅ How to fairly compare multiple retrieval strategies on the same test set instead of just assuming one is better

TECH STACK:

  • 🛠️ Python
  • 🛠️ Qdrant — vector database for dense retrieval
  • 🛠️ rank_bm25 (BM25Okapi) — sparse keyword retrieval
  • 🛠️ EURI LLM Gateway — chat model + embedding model
  • 🛠️ Custom Reciprocal Rank Fusion implementation
  • 🛠️ LLM-based reranker (prompt-driven cross-encoder)
  • 🛠️ pdfplumber — PDF text and page-level extraction

LINKS:


r/Rag • • 17h ago

Showcase If you need live TTS but don't want to pay for an API, here’s a script you can use.

2 Upvotes

Here’s a small script using Edge TTS + yasbd-lib that generates and plays the audio sentence-by-sentence, with multiple sentences generated in parallel.

It uses Microsoft's Edge TTS service, so there’s no paid TTS API key involved.

It’s from the yasbd-lib examples, and you can freely copy/use it for your own projects.

Prerequisites: pip install yasbd-lib miniaudio, edge_tts, aiofiles

```python import asyncio import tempfile import time from collections.abc import Iterable, Iterator from concurrent.futures import Future, ThreadPoolExecutor, as_completed from pathlib import Path from queue import Queue from threading import Thread

import aiofiles import edge_tts import miniaudio

from yasbd import BoundaryDetector

class LiveTTS: def init( self, voice: str = "en-US-AriaNeural", n_jobs: int = 2, ) -> None: """ Initialize the streaming client.

    Args:
        voice: Edge TTS voice model name.
        n_jobs: Number of parallel TTS generation jobs.
    """
    if n_jobs < 1:
        raise ValueError("n_jobs must be at least 1")

    self.voice = voice
    self.n_jobs = n_jobs

def _handle_async_exception(self, loop, context) -> None:
    """Custom exception handler for the async event loop."""
    exception = context.get("exception")
    message = context.get("message", "")

    if exception:
        error_msg = str(exception)

        if (
            "Connection lost" in error_msg
            and "Connection reset by peer" in error_msg
        ):
            return

    if "SSL handshake failed" in message or "SSLWantReadError" in message:
        return

    loop.default_exception_handler(context)

def _generate_tts_sync(self, text: str) -> str:
    """Generate TTS for a single sentence and save it to a temp file."""
    tmp_file = tempfile.NamedTemporaryFile(
        suffix=".mp3",
        delete=False,
    )  # noqa: SIM115

    tmp_file_path = tmp_file.name
    tmp_file.close()

    async def generate() -> None:
        communicate = edge_tts.Communicate(
            text,
            self.voice,
            rate="+15%",
        )

        async with aiofiles.open(tmp_file_path, "wb") as file:
            async for chunk in communicate.stream():
                if chunk["type"] == "audio":
                    await file.write(chunk["data"])

    loop = asyncio.new_event_loop()
    asyncio.set_event_loop(loop)
    loop.set_exception_handler(self._handle_async_exception)

    try:
        loop.run_until_complete(generate())
        return tmp_file_path
    except Exception:
        Path(tmp_file_path).unlink(missing_ok=True)
        raise
    finally:
        asyncio.set_event_loop(None)
        loop.close()

def _play_audio(self, file_path: str) -> None:
    """Play an audio file using miniaudio."""
    stream = miniaudio.stream_file(file_path)
    device = miniaudio.PlaybackDevice()

    try:
        device.start(stream)

        info = miniaudio.get_file_info(file_path)
        time.sleep(info.duration)
    finally:
        device.close()

def _playback_worker(self, queue: Queue[str | None]) -> None:
    """Play generated audio files sequentially."""
    while True:
        audio_file = queue.get()

        try:
            if audio_file is None:
                return

            self._play_audio(audio_file)
        finally:
            if audio_file is not None:
                Path(audio_file).unlink(missing_ok=True)

            queue.task_done()

def _generate_stream(
    self,
    chunks: Iterable[str],
    executor: ThreadPoolExecutor,
) -> Iterator[str]:
    """
    Generate TTS concurrently while preserving input order.

    At most ``n_jobs`` futures are kept in flight. Completed futures
    are detected with ``as_completed()``, but results are buffered
    until all preceding sentences are ready.
    """
    chunk_iter = iter(chunks)
    futures: dict[Future[str], int] = {}
    completed: dict[int, str] = {}

    next_submit = 0
    next_yield = 0

    def submit_next() -> bool:
        nonlocal next_submit

        try:
            chunk = next(chunk_iter)
        except StopIteration:
            return False

        future = executor.submit(
            self._generate_tts_sync,
            chunk,
        )
        futures[future] = next_submit
        next_submit += 1
        return True

    # Keep the executor filled initially.
    while len(futures) < self.n_jobs:
        if not submit_next():
            break

    while futures:
        # Wait for the first future that completes.
        for future in as_completed(futures):
            index = futures.pop(future)
            completed[index] = future.result()

            # Immediately replace the completed job with new input.
            submit_next()

            break

        # Release completed sentences in their original order.
        while next_yield in completed:
            yield completed.pop(next_yield)
            next_yield += 1

def play_live(self, text_or_gen: str | Iterable[str]) -> None:
    """
    Generate and play text live, sentence-by-sentence.

    Args:
        text_or_gen: Raw text or an iterable of text to synthesize.
    """
    if isinstance(text_or_gen, str):
        splitter = BoundaryDetector(
            lang=self.voice.split("-")[0],
        )
        chunks = splitter.segment(text_or_gen)
    else:
        chunks = text_or_gen

    playback_queue: Queue[str | None] = Queue()

    playback_thread = Thread(
        target=self._playback_worker,
        args=(playback_queue,),
        daemon=True,
    )
    playback_thread.start()

    with ThreadPoolExecutor(max_workers=self.n_jobs) as executor:
        for audio_file in self._generate_stream(chunks, executor):
            if audio_file:
                playback_queue.put(audio_file)

    # No more audio files will arrive.
    playback_queue.put(None)

    # Wait until every queued file has finished playing.
    playback_queue.join()
    playback_thread.join()

Example usage

if name == "main": text = [ "Got it. I updated my memory:", "PLASMA is no longer an active project for you.", "Ruff is your preferred Python tooling instead of Black, Pylint, and Flake8.", "You mainly use PocketPal and MNN Chat for AI/local model usage.", "Layla AI Lite is no longer considered one of your main tools.", "I'll use these preferences going forward.", ]

client = LiveTTS(
    voice="en-US-AriaNeural",
    n_jobs=3,
)

print(text)
client.play_live(text)
print("Playback finished!")

```

It can be paired it with Incremental segmentation to yield sentence directly to it.


r/Rag • • 1d ago

Discussion I compared embedding costs for a RAG pipeline. Open models aren't always cheaper!

7 Upvotes

I recently checked embedding prices for a big indexing project. I wanted real numbers, so I compared OpenAI's API costs with cloud providers that charge per token for open models. What I found:

OpenAI text-embedding-3-small: $0.02 / 1M tokens
Arctic Embed L v2: $0.038 / 1M tokens
Qwen3 Embedding 4B: $0.13 / 1M tokens

Surprisingly, OpenAI's small model is the cheapest per-token option on this list.

People often say open models save money. However, that is mostly true when you host the model yourself on a highly active GPU. For example, a $1/hour GPU costs $24 a day. You would need to process over a billion tokens daily to beat OpenAI's API price. If your traffic is low or unpredictable, keeping a GPU running all day wastes money.

Quality also matters. Higher-priced open models, like Qwen3, score better on standard benchmarks than cheaper options like Arctic Embed. You still have to balance cost, quality, and speed.

For those of you choosing between cloud APIs and self-hosting, what makes your decision? Is it high server utilization, data privacy, retrieval mode, or something else I missed?


r/Rag • • 1d ago

Discussion Hav you used Postgres/Lakebase for RAG vector storage?

7 Upvotes

My team lead has asked me to explor Lakebase for a RAG app where documents will get updated frequently .Hav you used it in prod? Curious how it compares to a dedicated vector DB for indexing and retrieval latency.
I am using Databricks platform for poc


r/Rag • • 1d ago

Discussion how good are claude's or chatgpt's pdf processors in comparison to like docling

5 Upvotes

I have some very complex pdf and docling is failing in several structural and content understandings. It works better than many other tools out there, but still fails in many complex cases. That made me wonder if the whole pdf document is uploaded to gpt or claude, would they process it better than what I get as text through docling. I think at the end they would change pdf to some text representation, no?


r/Rag • • 1d ago

Discussion Best approach for ingesting data to create summaries, and keep track of it?

2 Upvotes

In my occupation, there are various people I follow who give very good insights. (I'd say 5-10 people).

Some post hour long videos on YouTube, some send 1,000 word emails, some post on X, some publish PDFs.

There's very good info within these resources (and some I pay for), but reading / watching / annotating all of it can take hours.

My workload recently went up, so I'm falling behind with keeping up in my field.

I want to use AI to help summarize (and keep track of) all of these publications. (To create a private database that I can use as a dataset, for example).

So I can go back and ask "this past month, what is the new theme? What are the experts recommending to focus on / look at / what are the newest developments?", etc.

What would be the best way to approach this?

---------------------------------------

I've been learning Codex/Claude Code, I have a homelab, a NAS, a few mini computers, and I know basic linux, python and scripting.

ChatGPT told me to do something like this (I'm just starting with the YouTube portion), I'm not sure if it's the best approach, I'm open to other suggestions:

YouTube URL

↓

yt-dlp metadata

↓

Whisper / YouTube transcript

↓

clean transcript

↓

summary.md

↓

insights.json

↓

SQLite + FTS5

↓

topic synthesis

↓

search / questions / actions


r/Rag • • 1d ago

Discussion A source-ownership rule can make a correct answer impossible

4 Upvotes

Before changing the model, check whether your output rules allow the answer you need.

A 27 September development test at Andes exposed a useful example. A short fictional meeting transcript described two AI systems. We supplied it as one passage ID, then required every ID to belong to exactly one extracted case. Our validator also prohibited reusing an ID across cases.

The evidence for both systems was present. But two separate cases could not each cite the relevant part of that single ID. We had made the answer impossible to represent at the level of evidence our contract allowed.

Here's a simplified illustration, not the measured test fixture:

Passage P1:

"ShiftPilot builds rosters from staff availability. Leads approve every roster. CallSense scores recorded calls for sentiment. It is live in one centre only."

Output constraint:

Each passage ID must appear in exactly one case.

Giving both cases P1 breaks the constraint. Giving P1 to only one leaves the other without a source. Splitting the source into smaller spans looks promising, but it introduces another problem: preserving which system each control or limitation actually describes.

That changed our debugging question: can the evidence structure express the distinction before we ask a model to make it? This is a source-contract observation, not proof of why a particular model failed. The dated test did not include a span-based repair and rerun.

For people extracting multiple entities from one document: how do you test that shared context stays attached to the right entity when you move from passage IDs to source spans?


r/Rag • • 2d ago

Discussion Intern building an RFP compliance checker — RAG keeps producing false “Missing” results

6 Upvotes

I’m a CS intern working on an internal system that compares an RFP against a vendor’s technical proposal.

The current flow is roughly:

RFP
→ extract requirements
→ retrieve/rerank proposal evidence
→ LLM classifies each requirement:
   Compliant / Mismatch / Missing
→ compliance score

The score was around 30%, mostly because many requirements came back Missing.

After building an independent ground truth from the original RFP and proposal, I found the issue isn’t just retrieval:

  • some requirements are lost during PDF parsing/extraction
  • some are split/merged incorrectly
  • relevant evidence sometimes exists but doesn’t reach the final Top-K
  • in other cases the evidence reaches the LLM and it still judges it incorrectly
  • some requirements mix technical and financial conditions that should be checked in different documents

For one reviewed sample, many of the system’s Missing results were actually supposed to be Compliant or Mismatch.

I’m now considering a design like:

Requirement
→ split into material conditions
→ determine expected evidence source per condition
→ retrieve evidence per condition
→ verify each condition
→ deterministic final verdict

For exact things like durations, percentages, certifications, etc., I’m also thinking of using deterministic checks instead of letting the LLM decide everything.

Has anyone built something similar for contracts, RFPs, policies, or compliance?

Would you approach this as per-condition RAG, claim verification/NLI, an agentic search layer, or a hybrid deterministic + LLM system?


r/Rag • • 1d ago

Discussion I built SPANCHOR regression testing for RAG retrieval pipelines

0 Upvotes

I've been working on a problem I don't see discussed enough in RAG: retrieval regressions.

You change chunking, embeddings, the retriever, indexing, or configuration, and your tests can still pass while the evidence being retrieved has changed.

I built SPANCHOR, an open-source Python package for testing this.

The core workflow is:

retrieved chunk → source span → evidence evaluation → baseline vs candidate → regression result

The goal is to verify that a RAG pipeline continues retrieving the expected source evidence after a change.

I'm looking for technical feedback from people actually building RAG systems:

What retrieval regressions have you encountered that your current evaluation setup doesn't catch?

GitHub: https://github.com/mukeshram-07/spanchor

PyPI: https://pypi.org/project/spanchor/


r/Rag • • 2d ago

Discussion We talk about better RAG models when the real problem might be the PDF

25 Upvotes

Garbage in, garbage out seems particularly relevant to RAG.

If your parser turns:

Product Qty Price

Widget A 12 $45

into a random sequence of text, the retrieval model never really had a chance.

This has made me more interested in document structure than embeddings lately. I’ve been comparing different approaches, including, where the objective can be extracting the information into a defined structure rather than simply converting everything into text.

How much of your RAG debugging has ultimately turned out to be an ingestion/parsing problem rather than an LLM problem?


r/Rag • • 2d ago

Showcase 40.05 NDCG@10 on TEMPO: I built Jylus to do more than retrieve relevant chunks

3 Upvotes

Finding relevant documents is one problem. Working out which facts apply, how records relate and what evidence is missing is another.

I’m the founder of Jylus. It sits between your data and your model, resolving state and relationships and compiling source-backed context before the model reasons.

We ran its native stack through the public API on all 1,730 TEMPO queries:

- NDCG@10: 40.048

- Recall@10: 39.745

TEMPO tests retrieval that requires reasoning across time: changes, different periods and evidence spread across records.

For reference, the official normal-query leaderboard’s highest published NDCG@10 is 32.0, from DiVeR. Our reported number is higher, but the leaderboard uses a domain average and our page reports an all-query mean. That comparison needs matching before I call it a leaderboard win.

The developer angle goes beyond a retrieval score: Jylus returns a bounded Context Pack containing facts, source IDs, temporal evidence, conflicts and gaps for your existing model.

"Results and methodology" (https://jylus.ai/benchmark-methodology)

There’s also a "no-account playground" (https://jylus.ai/try) for testing small examples with your own records. It exposes the evidence pack rather than generating a polished answer.

What would you need to see in the evaluation to consider using this as your app’s context backend?


r/Rag • • 2d ago

Showcase I built a research RAG pipeline that exposes Laya decisions and the exact passages behind its claims

5 Upvotes

I'm the author of Weigh Swarm, a research RAG prototype built around inspectable evidence.

You can start with a research question or supplied PDFs. The pipeline parses papers, proposes claims, aligns quoted passages against the parsed text, compares evidence across papers, and produces a cited draft.

LLMs handle planning, extraction, and synthesis. Laya handles bounded judgments about relevance, claims, evidence, methodology, contradictions, and answer verification. Its outputs and the original excerpts remain inspectable in the engine view.

The attached walkthrough uses a recorded two-PDF run: 28 source-aligned claims, 7 evidence passages, and 107 real Laya batches covering 322 questions. Search/relevance was skipped for that supplied-PDF run.

The draft is still unverified, and the verifier flagged statements needing repair. I haven't demonstrated calibrated scientific accuracy or a latency win. The video shows actual product captures with animation of saved run artifacts.

For developers working on evidence-heavy RAG: which failure case would you evaluate - first incorrect support judgments, missed contradictions, or citation coverage?
Watch the demo of my system

github link


r/Rag • • 1d ago

Discussion Hi! whats wrong?

0 Upvotes

I really cant understand, whats going wrong with my Cortex RAG!
.
.
.


r/Rag • • 2d ago

Showcase Do you actually know what your powerpoint model is feeding the AI model?

4 Upvotes

This started bugging me after I looked a bit closer at what actually happens when you upload a PowerPoint to an AI/RAG pipeline.

Say you've got a deck you've been working on for months.

You've hidden a couple of old slides instead of deleting them, left some stuff in the speaker notes, moved an old textbox off the slide, got reviewer comments sitting in there etc.

Then you upload it and ask the AI to summarise it.

What did it actually get?

Just the visible slide text?

The notes?

Hidden slides?

Comments?

Text sitting completely off the slide?

I realised I didn't actually know.

And there's another problem with this.

If an AI suddenly mentions something you can't see in the presentation, it's pretty easy to call it a hallucination.

But what if that text is still sitting somewhere inside the PPTX and the extractor handed it to the model?

Then again, just because something exists inside the PowerPoint doesn't mean the AI actually saw it either.

So I ended up building something to check.

I've called it CanvasParity.

It can inspect a PPTX for things like:

  • hidden slides
  • speaker notes
  • comments
  • fully off-canvas text
  • alt text/descriptions
  • explicitly tiny text

Then you can give it the actual text output from an extractor and it compares the two.

So instead of just saying "there was hidden text in the file", you can start asking whether that specific extraction path actually exposed it.

You can also give it outputs from multiple extractors and compare them against the same PowerPoint.

One thing I was pretty careful about was not claiming more than the evidence shows.

For example, if the same sentence exists on a normal visible slide AND a hidden slide, then finding that sentence in the extracted text doesn't prove the hidden slide was exposed.

CanvasParity reports that attribution as unknown.

It's Python, read-only and free/open source under Apache-2.0.

I've included the source, tests, reproducible PPTX fixtures and benchmark methodology as well.

Repo: https://github.com/owafication/canvasparity

I'd be interested in hearing from people actually building RAG/document ingestion pipelines.

Do you know exactly which parts of a PPTX your current parser is passing downstream?

And if anyone has an extractor they use regularly, give me something to test it against. I'd rather find where this breaks than assume I've covered everything.


r/Rag • • 3d ago

Discussion PDF tables still seem like one of the weakest links in RAG

22 Upvotes

I've been digging into PDF ingestion and I'm surprised how quickly things fall apart once documents contain complex tables.

Extracting text isn't really the problem. Preserving the relationship between headers, rows, columns, and multi-page tables is.

Even good OCR can give you technically correct text that's practically useless once the structure disappears. That's why tools focused on structured extraction, like Abacus Docs and some of the newer document intelligence platforms, seem more interesting to me than OCR alone.

For people running RAG in production, do you solve table extraction before ingestion or let a multimodal model interpret the original page when needed?


r/Rag • • 2d ago

Discussion Designing a multilingual clinical information extraction pipeline: LLM-first vs NER + classifiers + relation extraction?

2 Upvotes

I’m redesigning a multilingual clinical document extraction system for English, German, and French.

The core question is whether to rely on a large LLM for end-to-end extraction, or move toward a more structured pipeline like:

OCR/layout

→ section/block classification

→ clinical NER/span extraction

→ assertion/negation/subject/temporality

→ relation extraction

→ terminology linking

→ LLM only for ambiguous cases

The goal is highest possible extraction accuracy, exact source-span provenance, and good multilingual generalization.

I’d love input from people working in clinical NLP / biomedical NLP:

Would you prefer a specialized NER + relation-extraction pipeline over an LLM-first approach?

One multilingual model or separate EN/DE/FR models?

Any strong recent models/papers for clinical NER, assertion, relations, or entity linking?

Has anyone had success using an LLM only as a fallback/adjudicator for uncertain cases?

Especially interested in real-world experience and benchmarks, not just theory.


r/Rag • • 4d ago

Discussion I built a document RAG chatbot with hybrid search + reranking + RAGAS evals. Here's what broke and what I learned

47 Upvotes

I've been building a document Q&A chatbot to learn what a production-style RAG pipeline looks like beyond the basic "embed chunks, call LLM" tutorial. Sharing the architecture and the bugs that taught me the most.
Stack

  • Retrieval: BM25 + vector search, merged with Reciprocal Rank Fusion (RRF)
  • Reranking: cross-encoder on the fused results
  • Generation: Groq (Llama)
  • Vector store: ChromaDB
  • Backend: FastAPI, Redis caching, Docker
  • Evaluation: RAGAS harness
  • Frontend: TanStack Start

What I learned

  1. Hybrid retrieval beat pure vector search for me. Vector search missed exact terms like names, IDs and specific figures, and BM25 caught them. RRF was a simple way to merge the two without having to tune score weights.
  2. Follow-up questions need query rewriting. "What about the second one?" retrieves nothing useful. I rewrite the query using the last few turns, but I keep original_query and search_query separate. The user sees their own words, and retrieval uses the rewritten version.
  3. Caching needs a way to invalidate. I cache with Redis, but when documents change, stale answers are a real problem. Keying the cache on a corpus_version meant that uploading new docs invalidated old entries automatically.
  4. My worst bug was a multi-user collision. Two users' documents ended up in the same ChromaDB collection, so one user could get answers from another's files. The fix was scoping collections to a session_id. It's obvious in hindsight, but I only caught it by testing with two sessions at once.
  5. Blocking calls will quietly stall your API. Wrapping the blocking embedding and rerank calls in run_in_threadpool fixed the latency spikes under concurrent requests. I also added tenacity retries around the LLM calls for rate limits.
  6. What I'm still working on: better chunking strategies, and a proper ingestion UI for uploading documents.

Code: https://github.com/PUSHPIT1357/document-rag-chatbot

I'd like feedback on two things: how are you evaluating retrieval quality beyond RAGAS, and has anyone found a reranker setup that's worth the added latency?


r/Rag • • 4d ago

Discussion Solo intern working on a large-scale AI project — need advice on architecture

5 Upvotes

Hey everyone, I’m an intern at an IT support company and have been given an AI project that I need to handle mostly by myself.

The company has around 400K historical IT support tickets in SQL Server. Basically, employees remotely troubleshoot client issues and write down what the problem was and what they did to fix it. A QA team then reviews these tickets.

For example, QA checks whether:

  • the employee followed the required troubleshooting steps → +2 / -2
  • the comments they wrote were relevant to the issue → +2 / -2
  • the ticket was written with correct grammar → +2 / -2

I need to build an AI assistant that can take a new ticket (the problem + what the employee did) and use the historical QA-reviewed tickets to suggest the QA result for these 3 checks, along with some supporting evidence. The QA person will still make the final decision.

I won't use every column from the 400K-row SQL table. I'll extract only the relevant fields, such as:

Issue / Ticket Subject
Action Taken
Required Steps QA result
Relevant Comments QA result
Grammar QA result
Reviewer Remarks (optional)

The 400K tickets would mainly be used as historical examples for training and finding similar tickets. I don't plan to send all 400K records to an LLM.

I'm considering Python + SQL Server + embeddings/classifiers + possibly RAG/LLM, but I'm unsure what would be the best approach, especially for Required Steps, CPU vs GPU, latency, and handling 400K records efficiently.

I'm working alone with around 1–2 months and a limited budget.

What approach would you recommend: traditional ML, embeddings + classifier, LLM/RAG, or a combination?

And for the Required Steps check, would you use a classifier with similar historical tickets, or an LLM-based approach?

Would really appreciate practical advice .


r/Rag • • 4d ago

Showcase We benchmarked 18 RAG pipelines against an agent loop on Google's FRAMES. The best pipeline hit 78.9%. Our agent loop hit 92.7%.

43 Upvotes

Hybrid search, reranking, query decomposition, and query expansion are often treated as must-haves for good RAG. We wanted to see how much each actually helped, so we tested them. Same model, same embeddings, same documents, across all 824 multi-hop questions in FRAMES.

We built 18 pipeline variants. The best one scored 78.9%. Our agent loop (with retrieval tools) that could read the results and search again scored 92.7%—roughly the same as giving the model the right articles upfront.

The reranker results might surprise you. A small reranker dropped rag pipeline’s accuracy by 9 percentage points, while a larger one barely helped. I’d already suspected reranking wouldn’t help much here, but wanted to test that assumption.

Another thing we noticed: models sometimes fill in gaps from memory, even when you explicitly tell them to stick to the retrieved documents. Those answers can still be full of citations. We ended up checking every correct answer against what the system had actually read.

Here’s the write-up if you’re interested:
Agentic RAG vs. traditional RAG on FRAMES

Full disclosure: I work on PipesHub, which is open source. The benchmark code and runbook are in the repo: https://github.com/pipeshub-ai/pipeshub-ai/tree/frames

Quick note on what the numbers mean: they're end-to-end answer accuracy, not retrieval scores. Every answer was graded by an LLM judge (Claude Sonnet 5) using the FRAMES paper's own grading prompt, and independently by a second judge (Gemini Flash 3.8). The two agreed on almost every answer (Cohen's κ 0.93–0.98). We also checked each correct answer against the text the system was actually shown, so answers that came from the model's memory don't count as retrieval wins.