CourseRAG · Module -1 :Foundations (What RAG Is and When It Is the Wrong Answer) · part 4 of 82
Part 4 · Module -1 :Foundations (What RAG Is and When It Is the Wrong Answer)

Topic 4: Failure Modes, Named Early

3 min read·21 Sept 2026

"The bot gave a bad answer" can't be fixed. "The right chunk ranked 14th and was cut off at top-5" can.

4.1 Retrieval vs Assembly vs Generation Failure

Intuition: A student fails an open-book exam because they opened the wrong page (retrieval), the right page was buried under other papers (assembly), or they read it and still wrote the wrong answer (generation).

Flowchart

Real story: a team tuned prompts for weeks. The real cause was that the warranty chunk was retrieved but cut off by a token limit, an assembly failure fixed with one line of code.

Watch out: log what was retrieved, what was reranked, what went into the prompt, and the answer. Without these four, you can't tell the failure types apart.

4.2 The Four Places the Answer Disappears

Check them in this order:

#FailureWhat happenedExampleFix
1Missing contentThe answer isn't in your documentsNo doc covers international shippingAdd content, or reply "I don't know" gracefully
2Extraction failureIt's in the file but lost in parsingRoaming rates were an image inside a PDFBetter parser, OCR
3Missed top-kIndexed, but ranked below the cutoffRight chunk at rank 12, k = 5Better chunking/embeddings, hybrid search
4Lost in rerankingRetrieved, then pushed out by the rerankerRank 3 → rank 9Better reranker; evaluate it on its own

Use cases

SituationMost likely failureUse thisWhy
Product launched yesterdayMissing contentFaster ingestionDocs aren't indexed yet
Price lists in scanned PDFsExtractionOCR + table extractionNo usable text layer
X200, X200 Pro, X300 manualsMissed top-kMetadata filter on model nameEmbeddings see them as nearly identical
General reranker on medical textLost in rerankingDomain-suited rerankerReranker doesn't understand the domain

4.3 Wrong-but-Fluent Answers: The Expensive Kind

Intuition: A GPS that says "signal lost" is annoying. A GPS that confidently drives you into a lake is dangerous.

FailureDo users notice?Cost
Error or crashImmediatelyLow
"I don't know"ImmediatelyLow to medium
Wrong but fluentOften neverHigh: bad decisions, refunds, legal risk, lost trust

How they happen: a similar-but-wrong chunk is retrieved, the model fills gaps from memory, old and new versions are retrieved together, or citations don't support the claims.

How to reduce them: reward "I don't know," verify citations in code, filter by product and version, check whether each claim appears in the context (faithfulness), and abstain when retrieval scores are low.

Real story: an airline's chatbot confidently described a bereavement refund rule that contradicted the airline's actual policy. A tribunal held the airline responsible for what its chatbot said.

4.4 The "It Works on My Ten Test Questions" Trap

Intuition: Testing a car only in your driveway doesn't prove it's safe on the highway.

Why ten questions mislead:

• Builders write questions their docs answer well.

• Test questions reuse document wording; real users don't ("return window" vs "how long till I can send it back").

• No typos, multi-part, or unanswerable questions.

• With 10 questions, one change swings the score by 10%, which is noise.

How to escape it:

1. Build 50-200+ test questions early and keep growing the set.

2. Add real user questions from logs.

3. Cover easy, paraphrased, multi-step, ambiguous, unanswerable, and typo questions.

4. Record the expected source chunk for each, so you can measure retrieval on its own.

5. Track recall@k (did the right chunk appear in the top k?) after every change.

Real story: a bot scored 10/10 in the demo. A 150-question set built from launch-week logs showed retrieval recall around 60%, and gave the team a real number to improve.