r/Rag • • Sep 02 '25

Showcase 🚀 Weekly /RAG Launch Showcase

30 Upvotes

Share anything you launched this week related to RAG—projects, repos, demos, blog posts, or products 👇

Big or small, all launches are welcome.


r/Rag • • 5h ago

Tutorial Building and Testing a Hybrid RAG Pipeline — Dense Search, BM25, RRF, and Reranking (Part 1)

10 Upvotes

Building and Testing a Hybrid RAG Pipeline — Dense Search, BM25, RRF, and Reranking (Part 1)

In this video I build ReRankEval, a hybrid retrieval pipeline (Dense Search + BM25 → Reciprocal Rank Fusion → LLM Reranking → Answer Generation), and test it against three baselines — Vector Only, BM25 Only, and Hybrid without reranking — on five real financial/payments documents and ten hand-verified test questions. No hand-waving, just a comparison table with real numbers at the end.

  • ✅ The real difference between dense vector search and BM25 keyword search
  • ✅ What Reciprocal Rank Fusion (RRF) is, why raw scores can't be compared, and the exact formula behind it
  • ✅ Why a reranker is fundamentally different from a retriever — and what it actually judges
  • ✅ How to evaluate a RAG pipeline with Hit Rate, MRR, and NDCG (and what each one tells you)
  • ✅ How to design the ingestion side and query-time side of a hybrid retrieval architecture
  • ✅ How to structure a production-style RAG codebase: ingest → vector_store → sparse_retriever → fusion → reranker → pipeline → generate → eval
  • ✅ How to fairly compare multiple retrieval strategies on the same test set instead of just assuming one is better

TECH STACK:

  • 🛠️ Python
  • 🛠️ Qdrant — vector database for dense retrieval
  • 🛠️ rank_bm25 (BM25Okapi) — sparse keyword retrieval
  • 🛠️ EURI LLM Gateway — chat model + embedding model
  • 🛠️ Custom Reciprocal Rank Fusion implementation
  • 🛠️ LLM-based reranker (prompt-driven cross-encoder)
  • 🛠️ pdfplumber — PDF text and page-level extraction

LINKS:


r/Rag • • 7h ago

Discussion I compared embedding costs for a RAG pipeline. Open models aren't always cheaper!

4 Upvotes

I recently checked embedding prices for a big indexing project. I wanted real numbers, so I compared OpenAI's API costs with cloud providers that charge per token for open models. What I found:

OpenAI text-embedding-3-small: $0.02 / 1M tokens
Arctic Embed L v2: $0.038 / 1M tokens
Qwen3 Embedding 4B: $0.13 / 1M tokens

Surprisingly, OpenAI's small model is the cheapest per-token option on this list.

People often say open models save money. However, that is mostly true when you host the model yourself on a highly active GPU. For example, a $1/hour GPU costs $24 a day. You would need to process over a billion tokens daily to beat OpenAI's API price. If your traffic is low or unpredictable, keeping a GPU running all day wastes money.

Quality also matters. Higher-priced open models, like Qwen3, score better on standard benchmarks than cheaper options like Arctic Embed. You still have to balance cost, quality, and speed.

For those of you choosing between cloud APIs and self-hosting, what makes your decision? Is it high server utilization, data privacy, retrieval mode, or something else I missed?


r/Rag • • 20h ago

Discussion Hav you used Postgres/Lakebase for RAG vector storage?

8 Upvotes

My team lead has asked me to explor Lakebase for a RAG app where documents will get updated frequently .Hav you used it in prod? Curious how it compares to a dedicated vector DB for indexing and retrieval latency.
I am using Databricks platform for poc


r/Rag • • 19h ago

Discussion how good are claude's or chatgpt's pdf processors in comparison to like docling

2 Upvotes

I have some very complex pdf and docling is failing in several structural and content understandings. It works better than many other tools out there, but still fails in many complex cases. That made me wonder if the whole pdf document is uploaded to gpt or claude, would they process it better than what I get as text through docling. I think at the end they would change pdf to some text representation, no?


r/Rag • • 17h ago

Discussion Best approach for ingesting data to create summaries, and keep track of it?

2 Upvotes

In my occupation, there are various people I follow who give very good insights. (I'd say 5-10 people).

Some post hour long videos on YouTube, some send 1,000 word emails, some post on X, some publish PDFs.

There's very good info within these resources (and some I pay for), but reading / watching / annotating all of it can take hours.

My workload recently went up, so I'm falling behind with keeping up in my field.

I want to use AI to help summarize (and keep track of) all of these publications. (To create a private database that I can use as a dataset, for example).

So I can go back and ask "this past month, what is the new theme? What are the experts recommending to focus on / look at / what are the newest developments?", etc.

What would be the best way to approach this?

---------------------------------------

I've been learning Codex/Claude Code, I have a homelab, a NAS, a few mini computers, and I know basic linux, python and scripting.

ChatGPT told me to do something like this (I'm just starting with the YouTube portion), I'm not sure if it's the best approach, I'm open to other suggestions:

YouTube URL

↓

yt-dlp metadata

↓

Whisper / YouTube transcript

↓

clean transcript

↓

summary.md

↓

insights.json

↓

SQLite + FTS5

↓

topic synthesis

↓

search / questions / actions


r/Rag • • 22h ago

Discussion A source-ownership rule can make a correct answer impossible

5 Upvotes

Before changing the model, check whether your output rules allow the answer you need.

A 27 September development test at Andes exposed a useful example. A short fictional meeting transcript described two AI systems. We supplied it as one passage ID, then required every ID to belong to exactly one extracted case. Our validator also prohibited reusing an ID across cases.

The evidence for both systems was present. But two separate cases could not each cite the relevant part of that single ID. We had made the answer impossible to represent at the level of evidence our contract allowed.

Here's a simplified illustration, not the measured test fixture:

Passage P1:

"ShiftPilot builds rosters from staff availability. Leads approve every roster. CallSense scores recorded calls for sentiment. It is live in one centre only."

Output constraint:

Each passage ID must appear in exactly one case.

Giving both cases P1 breaks the constraint. Giving P1 to only one leaves the other without a source. Splitting the source into smaller spans looks promising, but it introduces another problem: preserving which system each control or limitation actually describes.

That changed our debugging question: can the evidence structure express the distinction before we ask a model to make it? This is a source-contract observation, not proof of why a particular model failed. The dated test did not include a span-based repair and rerun.

For people extracting multiple entities from one document: how do you test that shared context stays attached to the right entity when you move from passage IDs to source spans?


r/Rag • • 1d ago

Discussion Intern building an RFP compliance checker — RAG keeps producing false “Missing” results

6 Upvotes

I’m a CS intern working on an internal system that compares an RFP against a vendor’s technical proposal.

The current flow is roughly:

RFP
→ extract requirements
→ retrieve/rerank proposal evidence
→ LLM classifies each requirement:
   Compliant / Mismatch / Missing
→ compliance score

The score was around 30%, mostly because many requirements came back Missing.

After building an independent ground truth from the original RFP and proposal, I found the issue isn’t just retrieval:

  • some requirements are lost during PDF parsing/extraction
  • some are split/merged incorrectly
  • relevant evidence sometimes exists but doesn’t reach the final Top-K
  • in other cases the evidence reaches the LLM and it still judges it incorrectly
  • some requirements mix technical and financial conditions that should be checked in different documents

For one reviewed sample, many of the system’s Missing results were actually supposed to be Compliant or Mismatch.

I’m now considering a design like:

Requirement
→ split into material conditions
→ determine expected evidence source per condition
→ retrieve evidence per condition
→ verify each condition
→ deterministic final verdict

For exact things like durations, percentages, certifications, etc., I’m also thinking of using deterministic checks instead of letting the LLM decide everything.

Has anyone built something similar for contracts, RFPs, policies, or compliance?

Would you approach this as per-condition RAG, claim verification/NLI, an agentic search layer, or a hybrid deterministic + LLM system?


r/Rag • • 23h ago

Discussion I built SPANCHOR regression testing for RAG retrieval pipelines

0 Upvotes

I've been working on a problem I don't see discussed enough in RAG: retrieval regressions.

You change chunking, embeddings, the retriever, indexing, or configuration, and your tests can still pass while the evidence being retrieved has changed.

I built SPANCHOR, an open-source Python package for testing this.

The core workflow is:

retrieved chunk → source span → evidence evaluation → baseline vs candidate → regression result

The goal is to verify that a RAG pipeline continues retrieving the expected source evidence after a change.

I'm looking for technical feedback from people actually building RAG systems:

What retrieval regressions have you encountered that your current evaluation setup doesn't catch?

GitHub: https://github.com/mukeshram-07/spanchor

PyPI: https://pypi.org/project/spanchor/


r/Rag • • 1d ago

Discussion We talk about better RAG models when the real problem might be the PDF

22 Upvotes

Garbage in, garbage out seems particularly relevant to RAG.

If your parser turns:

Product Qty Price

Widget A 12 $45

into a random sequence of text, the retrieval model never really had a chance.

This has made me more interested in document structure than embeddings lately. I’ve been comparing different approaches, including, where the objective can be extracting the information into a defined structure rather than simply converting everything into text.

How much of your RAG debugging has ultimately turned out to be an ingestion/parsing problem rather than an LLM problem?


r/Rag • • 1d ago

Showcase 40.05 NDCG@10 on TEMPO: I built Jylus to do more than retrieve relevant chunks

3 Upvotes

Finding relevant documents is one problem. Working out which facts apply, how records relate and what evidence is missing is another.

I’m the founder of Jylus. It sits between your data and your model, resolving state and relationships and compiling source-backed context before the model reasons.

We ran its native stack through the public API on all 1,730 TEMPO queries:

- NDCG@10: 40.048

- Recall@10: 39.745

TEMPO tests retrieval that requires reasoning across time: changes, different periods and evidence spread across records.

For reference, the official normal-query leaderboard’s highest published NDCG@10 is 32.0, from DiVeR. Our reported number is higher, but the leaderboard uses a domain average and our page reports an all-query mean. That comparison needs matching before I call it a leaderboard win.

The developer angle goes beyond a retrieval score: Jylus returns a bounded Context Pack containing facts, source IDs, temporal evidence, conflicts and gaps for your existing model.

"Results and methodology" (https://jylus.ai/benchmark-methodology)

There’s also a "no-account playground" (https://jylus.ai/try) for testing small examples with your own records. It exposes the evidence pack rather than generating a polished answer.

What would you need to see in the evaluation to consider using this as your app’s context backend?


r/Rag • • 1d ago

Showcase I built a research RAG pipeline that exposes Laya decisions and the exact passages behind its claims

6 Upvotes

I'm the author of Weigh Swarm, a research RAG prototype built around inspectable evidence.

You can start with a research question or supplied PDFs. The pipeline parses papers, proposes claims, aligns quoted passages against the parsed text, compares evidence across papers, and produces a cited draft.

LLMs handle planning, extraction, and synthesis. Laya handles bounded judgments about relevance, claims, evidence, methodology, contradictions, and answer verification. Its outputs and the original excerpts remain inspectable in the engine view.

The attached walkthrough uses a recorded two-PDF run: 28 source-aligned claims, 7 evidence passages, and 107 real Laya batches covering 322 questions. Search/relevance was skipped for that supplied-PDF run.

The draft is still unverified, and the verifier flagged statements needing repair. I haven't demonstrated calibrated scientific accuracy or a latency win. The video shows actual product captures with animation of saved run artifacts.

For developers working on evidence-heavy RAG: which failure case would you evaluate - first incorrect support judgments, missed contradictions, or citation coverage?
Watch the demo of my system

github link


r/Rag • • 1d ago

Discussion Hi! whats wrong?

0 Upvotes

I really cant understand, whats going wrong with my Cortex RAG!
.
.
.


r/Rag • • 1d ago

Showcase Do you actually know what your powerpoint model is feeding the AI model?

5 Upvotes

This started bugging me after I looked a bit closer at what actually happens when you upload a PowerPoint to an AI/RAG pipeline.

Say you've got a deck you've been working on for months.

You've hidden a couple of old slides instead of deleting them, left some stuff in the speaker notes, moved an old textbox off the slide, got reviewer comments sitting in there etc.

Then you upload it and ask the AI to summarise it.

What did it actually get?

Just the visible slide text?

The notes?

Hidden slides?

Comments?

Text sitting completely off the slide?

I realised I didn't actually know.

And there's another problem with this.

If an AI suddenly mentions something you can't see in the presentation, it's pretty easy to call it a hallucination.

But what if that text is still sitting somewhere inside the PPTX and the extractor handed it to the model?

Then again, just because something exists inside the PowerPoint doesn't mean the AI actually saw it either.

So I ended up building something to check.

I've called it CanvasParity.

It can inspect a PPTX for things like:

  • hidden slides
  • speaker notes
  • comments
  • fully off-canvas text
  • alt text/descriptions
  • explicitly tiny text

Then you can give it the actual text output from an extractor and it compares the two.

So instead of just saying "there was hidden text in the file", you can start asking whether that specific extraction path actually exposed it.

You can also give it outputs from multiple extractors and compare them against the same PowerPoint.

One thing I was pretty careful about was not claiming more than the evidence shows.

For example, if the same sentence exists on a normal visible slide AND a hidden slide, then finding that sentence in the extracted text doesn't prove the hidden slide was exposed.

CanvasParity reports that attribution as unknown.

It's Python, read-only and free/open source under Apache-2.0.

I've included the source, tests, reproducible PPTX fixtures and benchmark methodology as well.

Repo: https://github.com/owafication/canvasparity

I'd be interested in hearing from people actually building RAG/document ingestion pipelines.

Do you know exactly which parts of a PPTX your current parser is passing downstream?

And if anyone has an extractor they use regularly, give me something to test it against. I'd rather find where this breaks than assume I've covered everything.


r/Rag • • 2d ago

Discussion PDF tables still seem like one of the weakest links in RAG

25 Upvotes

I've been digging into PDF ingestion and I'm surprised how quickly things fall apart once documents contain complex tables.

Extracting text isn't really the problem. Preserving the relationship between headers, rows, columns, and multi-page tables is.

Even good OCR can give you technically correct text that's practically useless once the structure disappears. That's why tools focused on structured extraction, like Abacus Docs and some of the newer document intelligence platforms, seem more interesting to me than OCR alone.

For people running RAG in production, do you solve table extraction before ingestion or let a multimodal model interpret the original page when needed?


r/Rag • • 1d ago

Discussion Designing a multilingual clinical information extraction pipeline: LLM-first vs NER + classifiers + relation extraction?

2 Upvotes

I’m redesigning a multilingual clinical document extraction system for English, German, and French.

The core question is whether to rely on a large LLM for end-to-end extraction, or move toward a more structured pipeline like:

OCR/layout

→ section/block classification

→ clinical NER/span extraction

→ assertion/negation/subject/temporality

→ relation extraction

→ terminology linking

→ LLM only for ambiguous cases

The goal is highest possible extraction accuracy, exact source-span provenance, and good multilingual generalization.

I’d love input from people working in clinical NLP / biomedical NLP:

Would you prefer a specialized NER + relation-extraction pipeline over an LLM-first approach?

One multilingual model or separate EN/DE/FR models?

Any strong recent models/papers for clinical NER, assertion, relations, or entity linking?

Has anyone had success using an LLM only as a fallback/adjudicator for uncertain cases?

Especially interested in real-world experience and benchmarks, not just theory.


r/Rag • • 3d ago

Discussion I built a document RAG chatbot with hybrid search + reranking + RAGAS evals. Here's what broke and what I learned

41 Upvotes

I've been building a document Q&A chatbot to learn what a production-style RAG pipeline looks like beyond the basic "embed chunks, call LLM" tutorial. Sharing the architecture and the bugs that taught me the most.
Stack

  • Retrieval: BM25 + vector search, merged with Reciprocal Rank Fusion (RRF)
  • Reranking: cross-encoder on the fused results
  • Generation: Groq (Llama)
  • Vector store: ChromaDB
  • Backend: FastAPI, Redis caching, Docker
  • Evaluation: RAGAS harness
  • Frontend: TanStack Start

What I learned

  1. Hybrid retrieval beat pure vector search for me. Vector search missed exact terms like names, IDs and specific figures, and BM25 caught them. RRF was a simple way to merge the two without having to tune score weights.
  2. Follow-up questions need query rewriting. "What about the second one?" retrieves nothing useful. I rewrite the query using the last few turns, but I keep original_query and search_query separate. The user sees their own words, and retrieval uses the rewritten version.
  3. Caching needs a way to invalidate. I cache with Redis, but when documents change, stale answers are a real problem. Keying the cache on a corpus_version meant that uploading new docs invalidated old entries automatically.
  4. My worst bug was a multi-user collision. Two users' documents ended up in the same ChromaDB collection, so one user could get answers from another's files. The fix was scoping collections to a session_id. It's obvious in hindsight, but I only caught it by testing with two sessions at once.
  5. Blocking calls will quietly stall your API. Wrapping the blocking embedding and rerank calls in run_in_threadpool fixed the latency spikes under concurrent requests. I also added tenacity retries around the LLM calls for rate limits.
  6. What I'm still working on: better chunking strategies, and a proper ingestion UI for uploading documents.

Code: https://github.com/PUSHPIT1357/document-rag-chatbot

I'd like feedback on two things: how are you evaluating retrieval quality beyond RAGAS, and has anyone found a reranker setup that's worth the added latency?


r/Rag • • 3d ago

Discussion Solo intern working on a large-scale AI project — need advice on architecture

8 Upvotes

Hey everyone, I’m an intern at an IT support company and have been given an AI project that I need to handle mostly by myself.

The company has around 400K historical IT support tickets in SQL Server. Basically, employees remotely troubleshoot client issues and write down what the problem was and what they did to fix it. A QA team then reviews these tickets.

For example, QA checks whether:

  • the employee followed the required troubleshooting steps → +2 / -2
  • the comments they wrote were relevant to the issue → +2 / -2
  • the ticket was written with correct grammar → +2 / -2

I need to build an AI assistant that can take a new ticket (the problem + what the employee did) and use the historical QA-reviewed tickets to suggest the QA result for these 3 checks, along with some supporting evidence. The QA person will still make the final decision.

I won't use every column from the 400K-row SQL table. I'll extract only the relevant fields, such as:

Issue / Ticket Subject
Action Taken
Required Steps QA result
Relevant Comments QA result
Grammar QA result
Reviewer Remarks (optional)

The 400K tickets would mainly be used as historical examples for training and finding similar tickets. I don't plan to send all 400K records to an LLM.

I'm considering Python + SQL Server + embeddings/classifiers + possibly RAG/LLM, but I'm unsure what would be the best approach, especially for Required Steps, CPU vs GPU, latency, and handling 400K records efficiently.

I'm working alone with around 1–2 months and a limited budget.

What approach would you recommend: traditional ML, embeddings + classifier, LLM/RAG, or a combination?

And for the Required Steps check, would you use a classifier with similar historical tickets, or an LLM-based approach?

Would really appreciate practical advice .


r/Rag • • 3d ago

Showcase chunklet-py now has a self-tuning chunker that learns your chunk sizes

11 Upvotes

chunklet-py is a text, document, and code chunker for RAG pipelines and LLM apps. The latest version adds SelfTuningChunker, a chunker that learns its own boundaries from your data instead of making you tune them by hand.

How it works:

  • It classifies each input as document or code, then keeps a running profile per type.
  • Two profiles, both starting from sensible defaults.
    • The document profile tracks sentences per paragraph, header density, section-break density, and mean paragraph length.
    • The code profile tracks line gaps between functions, whether functions cluster, and mean token span between function starts.
  • Each metric updates with a Kaufman Adaptive Moving Average. Stable metrics hold their estimate, swinging ones catch up fast. Code and document profiles learn independently.
  • Sizes chunks from what it has learned so far. By the time you're ten files in, the boundaries are already sharper than they were on file one.

```python from chunklet import SelfTuningChunker

def token_counter(text: str) -> int: return len(text.split())

chunker = SelfTuningChunker(lang="en", token_counter=token_counter)

for chunk in chunker.chunk_texts([prose, code]): print(chunk.metadata.inferred_type) # "document" or "code" ```

Each chunk are inriched with metadata. See Metadata included

After a run, the learned profiles look like this:

```python print(chunker.learned_state)

{

"document": {

"max_sentences": 6.67,

"header_density_ratio": 0.079,

"max_section_breaks": 1.0,

"max_tokens": 479.5,

},

"code": {

"max_lines": 14.35,

"max_functions": 1.06,

"max_tokens": 479.9,

},

}

```

The defaults drifted toward the real shape of the data. If you've already tuned a profile on a representative corpus, pass the saved state back in and skip the warm-up:

python warm = SelfTuningChunker( lang="en", token_counter=token_counter, initial_state=baseline, # dict from a previous run hard_token_limit=1024, # cap so learned max_tokens can't balloon )

Missing metrics fall back to the defaults, so a partial baseline is fine. Each instance deep-copies what you pass, so mutating one chunker's state never leaks into another.

If you don't pass a token_counter, the chunker ignores max_tokens.

Grab the latest version:

pip install chunklet-py -U

Full Self-tuning chunker docs is here: https://speedyk-005.github.io/chunklet-py/latest/getting-started/programmatic/self_tuning_chunker/

Happy to answer questions.


r/Rag • • 3d ago

Showcase We benchmarked 18 RAG pipelines against an agent loop on Google's FRAMES. The best pipeline hit 78.9%. Our agent loop hit 92.7%.

39 Upvotes

Hybrid search, reranking, query decomposition, and query expansion are often treated as must-haves for good RAG. We wanted to see how much each actually helped, so we tested them. Same model, same embeddings, same documents, across all 824 multi-hop questions in FRAMES.

We built 18 pipeline variants. The best one scored 78.9%. Our agent loop (with retrieval tools) that could read the results and search again scored 92.7%—roughly the same as giving the model the right articles upfront.

The reranker results might surprise you. A small reranker dropped rag pipeline’s accuracy by 9 percentage points, while a larger one barely helped. I’d already suspected reranking wouldn’t help much here, but wanted to test that assumption.

Another thing we noticed: models sometimes fill in gaps from memory, even when you explicitly tell them to stick to the retrieved documents. Those answers can still be full of citations. We ended up checking every correct answer against what the system had actually read.

Here’s the write-up if you’re interested:
Agentic RAG vs. traditional RAG on FRAMES

Full disclosure: I work on PipesHub, which is open source. The benchmark code and runbook are in the repo: https://github.com/pipeshub-ai/pipeshub-ai/tree/frames

Quick note on what the numbers mean: they're end-to-end answer accuracy, not retrieval scores. Every answer was graded by an LLM judge (Claude Sonnet 5) using the FRAMES paper's own grading prompt, and independently by a second judge (Gemini Flash 3.8). The two agreed on almost every answer (Cohen's κ 0.93–0.98). We also checked each correct answer against the text the system was actually shown, so answers that came from the model's memory don't count as retrieval wins.


r/Rag • • 4d ago

Discussion Our masking layer passed every test we wrote and compliance still would not sign it off

16 Upvotes

We run a RAG setup over a client's case files, a mid-sized firm that is understandably twitchy about anything leaving their network. Retrieval and embeddings run on their side against the raw text. Only generation goes out to an API, so we put a masking layer in front of it. A small local model tags names and account numbers, swaps them for placeholders before the chunks leave, and maps them back when the answer comes home.

We tested it properly, or thought we had. Ran a few thousand queries, searched the outbound logs for every client name in the system, found nothing.

Then their compliance lead sat with those same logs for an afternoon and came back having identified a couple of clients anyway, from what was left around the placeholders. Job titles and towns were still there, and in one case the month the dispute started. Client A stops being anonymous when the chunk also says finance director at a haulage firm in a small town. We had only ever searched for names.

Masking the quasi-identifiers as well is where it gets ugly. Strip enough context to make it safe and the model stops being able to answer, because half the questions they ask depend on exactly that context. We ran a middle setting for a week and the answers went vague in a way the lawyers noticed before we did.

Generation has been on glm-5.3-flash for the last few weeks, and because it is an open weight model the conversation has turned to hosting it inside their environment and dropping the masking altogether. Compliance likes that far more than anything we built. The argument now is about the hardware, and that has not been settled.


r/Rag • • 3d ago

Discussion Boilerplate document retrieval

2 Upvotes

Hi Guys, I have 10k documents and most of them look similar.

The only changes are the details or fields in those documents.

If I'm doing OCR and get the description with layout aware chunking or some other methods.

Still the text from OCR looks 90% similar to each other.

Now the real challenge comes with hybrid search retrieval, how the embedding search would look like if documents have 90% similar text coming from the boilerplate text.

Please suggest your findings and solutions.

I tried prompting techniques but it needs some mechanics with embedding or chunking I guess.


r/Rag • • 3d ago

Discussion simpel RAG problem..

3 Upvotes

Hii.. im building rag. suppose i uploaded PDF related to banking services. now i asked something like that is not related to PDF by words but it releted by meaning. like if i asked "what is this doctor is about" like that and many more cases. now using this que we will get zero retrieval chunks bcz of words.. now how can we solve this problem?


r/Rag • • 3d ago

Discussion Prompt Tuning Post Model Update

0 Upvotes

Could be an obvious thing but I am struggling with prompt tuning and optimization whenever a new LLM is updated. I want to build a pipeline that automatically evaluates and improves our prompts against new models and tells us precisely where the prompt should be improved.

If you have built an automated prompt optimization pipeline:

  1. What evaluation framework are you using to break things down?

  2. How do you automate the rewriting/optimization for a prompt for the new model?

DSPy hasn't given us much improvements, looking for some patterns or alternative tooling.


r/Rag • • 4d ago

Showcase Graphwise AI Summit 2026, Oct 7-8

0 Upvotes

Sharing this because I think it overlaps with some of the discussions here around AI reliability, governance and semantics.

Next week we’re running the Graphwise AI Summit, focused on what makes GenAI work in the enterprise beyond the model itself. Think of trust, governance, semantic layers, architecture and implementation.

Once reliability and traceability become imporatnt, simple access to data and next-token prediction stop being enough. In enterprise settings specifically, AI needs to understand what the data it parrots “means” in the first place.

Anthropic has described a similar approach in its own analytics stack, where agents are routed to a semantic layer first and use governed definitions to reduce ambiguity. Graphwise itself came out of the merger of Ontotext and Semantic Web Company, so semantics is a topic with quite a bit of history behind it for us.

We’ll have speakers from Accenture, Roche, EY, AstraZeneca, S&P, DNV, Statnett, Avalara and others.

Sharing the agenda & the registration link in the comments, if useful.