Topic 4: Failure Modes, Named Early
"The bot gave a bad answer" can't be fixed. "The right chunk ranked 14th and was cut off at top-5" can.
4.1 Retrieval vs Assembly vs Generation Failure
Intuition: A student fails an open-book exam because they opened the wrong page (retrieval), the right page was buried under other papers (assembly), or they read it and still wrote the wrong answer (generation).
Real story: a team tuned prompts for weeks. The real cause was that the warranty chunk was retrieved but cut off by a token limit, an assembly failure fixed with one line of code.
Watch out: log what was retrieved, what was reranked, what went into the prompt, and the answer. Without these four, you can't tell the failure types apart.
4.2 The Four Places the Answer Disappears
Check them in this order:
| # | Failure | What happened | Example | Fix |
| 1 | Missing content | The answer isn't in your documents | No doc covers international shipping | Add content, or reply "I don't know" gracefully |
| 2 | Extraction failure | It's in the file but lost in parsing | Roaming rates were an image inside a PDF | Better parser, OCR |
| 3 | Missed top-k | Indexed, but ranked below the cutoff | Right chunk at rank 12, k = 5 | Better chunking/embeddings, hybrid search |
| 4 | Lost in reranking | Retrieved, then pushed out by the reranker | Rank 3 → rank 9 | Better reranker; evaluate it on its own |
Use cases
| Situation | Most likely failure | Use this | Why |
| Product launched yesterday | Missing content | Faster ingestion | Docs aren't indexed yet |
| Price lists in scanned PDFs | Extraction | OCR + table extraction | No usable text layer |
| X200, X200 Pro, X300 manuals | Missed top-k | Metadata filter on model name | Embeddings see them as nearly identical |
| General reranker on medical text | Lost in reranking | Domain-suited reranker | Reranker doesn't understand the domain |
4.3 Wrong-but-Fluent Answers: The Expensive Kind
Intuition: A GPS that says "signal lost" is annoying. A GPS that confidently drives you into a lake is dangerous.
| Failure | Do users notice? | Cost |
| Error or crash | Immediately | Low |
| "I don't know" | Immediately | Low to medium |
| Wrong but fluent | Often never | High: bad decisions, refunds, legal risk, lost trust |
How they happen: a similar-but-wrong chunk is retrieved, the model fills gaps from memory, old and new versions are retrieved together, or citations don't support the claims.
How to reduce them: reward "I don't know," verify citations in code, filter by product and version, check whether each claim appears in the context (faithfulness), and abstain when retrieval scores are low.
Real story: an airline's chatbot confidently described a bereavement refund rule that contradicted the airline's actual policy. A tribunal held the airline responsible for what its chatbot said.
4.4 The "It Works on My Ten Test Questions" Trap
Intuition: Testing a car only in your driveway doesn't prove it's safe on the highway.
Why ten questions mislead:
• Builders write questions their docs answer well.
• Test questions reuse document wording; real users don't ("return window" vs "how long till I can send it back").
• No typos, multi-part, or unanswerable questions.
• With 10 questions, one change swings the score by 10%, which is noise.
How to escape it:
1. Build 50-200+ test questions early and keep growing the set.
2. Add real user questions from logs.
3. Cover easy, paraphrased, multi-step, ambiguous, unanswerable, and typo questions.
4. Record the expected source chunk for each, so you can measure retrieval on its own.
5. Track recall@k (did the right chunk appear in the top k?) after every change.
Real story: a bot scored 10/10 in the demo. A 150-question set built from launch-week logs showed retrieval recall around 60%, and gave the team a real number to improve.