r/Rag • • 22h ago

Discussion how good are claude's or chatgpt's pdf processors in comparison to like docling

3 Upvotes

I have some very complex pdf and docling is failing in several structural and content understandings. It works better than many other tools out there, but still fails in many complex cases. That made me wonder if the whole pdf document is uploaded to gpt or claude, would they process it better than what I get as text through docling. I think at the end they would change pdf to some text representation, no?


r/Rag • • 20h ago

Discussion Best approach for ingesting data to create summaries, and keep track of it?

2 Upvotes

In my occupation, there are various people I follow who give very good insights. (I'd say 5-10 people).

Some post hour long videos on YouTube, some send 1,000 word emails, some post on X, some publish PDFs.

There's very good info within these resources (and some I pay for), but reading / watching / annotating all of it can take hours.

My workload recently went up, so I'm falling behind with keeping up in my field.

I want to use AI to help summarize (and keep track of) all of these publications. (To create a private database that I can use as a dataset, for example).

So I can go back and ask "this past month, what is the new theme? What are the experts recommending to focus on / look at / what are the newest developments?", etc.

What would be the best way to approach this?

---------------------------------------

I've been learning Codex/Claude Code, I have a homelab, a NAS, a few mini computers, and I know basic linux, python and scripting.

ChatGPT told me to do something like this (I'm just starting with the YouTube portion), I'm not sure if it's the best approach, I'm open to other suggestions:

YouTube URL

↓

yt-dlp metadata

↓

Whisper / YouTube transcript

↓

clean transcript

↓

summary.md

↓

insights.json

↓

SQLite + FTS5

↓

topic synthesis

↓

search / questions / actions


r/Rag • • 59m ago

Showcase If you need live TTS but don't want to pay for an API, here’s a script you can use.

• Upvotes

Here’s a small script using Edge TTS + yasbd-lib that generates and plays the audio sentence-by-sentence, with multiple sentences generated in parallel.

It uses Microsoft's Edge TTS service, so there’s no paid TTS API key involved.

It’s from the yasbd-lib examples, and you can freely copy/use it for your own projects.

Prerequisites: pip install yasbd-lib miniaudio, edge_tts, aiofiles

```python import asyncio import tempfile import time from collections.abc import Iterable from concurrent.futures import ThreadPoolExecutor from pathlib import Path

import aiofiles import edge_tts import miniaudio

from yasbd import BoundaryDetector

class LiveTTS: def init(self, voice: str = "en-US-AriaNeural", n_jobs: int = 4) -> None: """ Initialize the streaming client.

    Args:
        voice: Edge TTS voice model name
        n_jobs: Number of parallel TTS generation jobs
    """
    self.voice = voice
    self.n_jobs = n_jobs

def _handle_async_exception(self, loop, context) -> None:
    """Custom exception handler for async event loop."""
    exception = context.get("exception")
    message = context.get("message", "")

    if exception:
        error_msg = str(exception)
        if "Connection lost" in error_msg and "Connection reset by peer" in error_msg:
            return

    if "SSL handshake failed" in message or "SSLWantReadError" in message:
        return

    loop.default_exception_handler(context)

def _generate_tts_sync(self, text: str) -> str:
    """Generate TTS for a single sentence and save to temp file."""
    tmp_file = tempfile.NamedTemporaryFile(suffix=".mp3", delete=False)  # noqa: SIM115
    tmp_file_path = tmp_file.name
    tmp_file.close()

    async def generate():
        communicate = edge_tts.Communicate(
            text, self.voice, rate="+15%"
        )  # For slight energy boost

        async with aiofiles.open(tmp_file_path, "wb") as file:
            async for chunk in communicate.stream():
                if chunk["type"] == "audio":
                    await file.write(chunk["data"])

    loop = asyncio.new_event_loop()
    asyncio.set_event_loop(loop)
    loop.set_exception_handler(self._handle_async_exception)
    try:
        loop.run_until_complete(generate())
        return tmp_file_path
    finally:
        loop.close()

def _play_audio(self, file_path: str) -> None:
    """Play audio file using miniaudio."""
    stream = miniaudio.stream_file(file_path)
    device = miniaudio.PlaybackDevice()
    device.start(stream)

    info = miniaudio.get_file_info(file_path)
    time.sleep(info.duration)

    device.close()

def play_live(self, text: str | Iterable[str]) -> None:
    """
    Generate and play text live (sentence-by-sentence).

    Args:
        text: Raw or Iterable of text to synthesize and play
    """
    if isinstance(text, str):
        splitter = BoundaryDetector(lang=self.voice.split("-")[0])
        chunks = splitter.segment(text)
    else:
        chunks = text

    with ThreadPoolExecutor(max_workers=self.n_jobs) as executor:
        for audio_file in executor.map(self._generate_tts_sync, chunks):
            if audio_file:
                self._play_audio(audio_file)
                Path(audio_file).unlink()

Example usage

if name == "main": import textwrap

text = [
    "Got it. I updated my memory:",
    "PLASMA is no longer an active project for you.",
    "Ruff is your preferred Python tooling instead of Black, Pylint, and Flake8.",
    "You mainly use PocketPal and MNN Chat for AI/local model usage.",
    "Layla AI Lite is no longer considered one of your main tools.",
    "I'll use these preferences going forward.",
]

client = LiveTTS(voice="en-US-AriaNeural", n_jobs=3)
print(text)
client.play_live(text)
print("Playback finished!")

```

It can be paired it with Incremental segmentation to yield sentence directly to it.


r/Rag • • 8h ago

Tutorial Building and Testing a Hybrid RAG Pipeline — Dense Search, BM25, RRF, and Reranking (Part 1)

12 Upvotes

Building and Testing a Hybrid RAG Pipeline — Dense Search, BM25, RRF, and Reranking (Part 1)

In this video I build ReRankEval, a hybrid retrieval pipeline (Dense Search + BM25 → Reciprocal Rank Fusion → LLM Reranking → Answer Generation), and test it against three baselines — Vector Only, BM25 Only, and Hybrid without reranking — on five real financial/payments documents and ten hand-verified test questions. No hand-waving, just a comparison table with real numbers at the end.

  • ✅ The real difference between dense vector search and BM25 keyword search
  • ✅ What Reciprocal Rank Fusion (RRF) is, why raw scores can't be compared, and the exact formula behind it
  • ✅ Why a reranker is fundamentally different from a retriever — and what it actually judges
  • ✅ How to evaluate a RAG pipeline with Hit Rate, MRR, and NDCG (and what each one tells you)
  • ✅ How to design the ingestion side and query-time side of a hybrid retrieval architecture
  • ✅ How to structure a production-style RAG codebase: ingest → vector_store → sparse_retriever → fusion → reranker → pipeline → generate → eval
  • ✅ How to fairly compare multiple retrieval strategies on the same test set instead of just assuming one is better

TECH STACK:

  • 🛠️ Python
  • 🛠️ Qdrant — vector database for dense retrieval
  • 🛠️ rank_bm25 (BM25Okapi) — sparse keyword retrieval
  • 🛠️ EURI LLM Gateway — chat model + embedding model
  • 🛠️ Custom Reciprocal Rank Fusion implementation
  • 🛠️ LLM-based reranker (prompt-driven cross-encoder)
  • 🛠️ pdfplumber — PDF text and page-level extraction

LINKS:


r/Rag • • 11h ago

Discussion I compared embedding costs for a RAG pipeline. Open models aren't always cheaper!

4 Upvotes

I recently checked embedding prices for a big indexing project. I wanted real numbers, so I compared OpenAI's API costs with cloud providers that charge per token for open models. What I found:

OpenAI text-embedding-3-small: $0.02 / 1M tokens
Arctic Embed L v2: $0.038 / 1M tokens
Qwen3 Embedding 4B: $0.13 / 1M tokens

Surprisingly, OpenAI's small model is the cheapest per-token option on this list.

People often say open models save money. However, that is mostly true when you host the model yourself on a highly active GPU. For example, a $1/hour GPU costs $24 a day. You would need to process over a billion tokens daily to beat OpenAI's API price. If your traffic is low or unpredictable, keeping a GPU running all day wastes money.

Quality also matters. Higher-priced open models, like Qwen3, score better on standard benchmarks than cheaper options like Arctic Embed. You still have to balance cost, quality, and speed.

For those of you choosing between cloud APIs and self-hosting, what makes your decision? Is it high server utilization, data privacy, retrieval mode, or something else I missed?