CourseRAG · Module -1 :Foundations (What RAG Is and When It Is the Wrong Answer) · part 1 of 82
Part 1 · Module -1 :Foundations (What RAG Is and When It Is the Wrong Answer)

Topic 1: Why RAG Exists (The Core Problem)

5 min read·21 Sept 2026

Module 1: Foundations (What RAG Is and When It Is the Wrong Answer)

By the end of this module, you'll be able to:
- Explain what problem RAG solves, and what it doesn't
- Draw the RAG pipeline from memory and say where each stage breaks
- Choose between RAG, long context, fine-tuning, tool calling, and agentic search, with reasons
- Name a failure precisely so you can fix it
- Set up the tech stack used for the rest of the course and build your first grounded RAG assistant

How this module is organized

TopicWhat you'll learnStyle
1. Why RAG ExistsThe problem RAG solvesConcepts
2. Anatomy of a RAG SystemThe pipeline and where it breaksConcepts
3. The Decision That Comes FirstWhen to use RAG, and when not toConcepts
4. Failure Modes, Named EarlyHow to diagnose bad answersConcepts
5. The Course Tech StackLLMs, embedding models, rerankers, vector DBs, frameworksSetup + examples
6. Hands-On LabBuild ShopSphere's first grounded assistantCode

Topics 1-4 are about thinking, so there's no code. You'll apply every idea in Topic 6.

The Running Project: ShopSphere Support Assistant

Throughout the course you'll build one assistant for ShopSphere, a fictional online store. It has four data sources:

Data sourceTypeSizeChanges
Help-center articles (returns, shipping)Text~2,000 pagesWeekly
Product manualsPDFs~15,000 pagesMonthly
Orders databaseTableMillions of rowsEvery second
Pricing APIAPIThousands of productsHourly

Keep these in mind. By the end of Topic 3, you'll know exactly which approach each one needs.

Topic 1: Why RAG Exists (The Core Problem)

1.1 Parametric Knowledge vs External Knowledge

Intuition: An LLM answering from memory is a student taking a closed-book exam. RAG turns it into an open-book exam: the student still needs to be smart, but now they can look things up.

Parametric knowledge is what the model learned in training, stored in its weights. It has four limits:

1. Frozen: it stops at the training date.

2. Public only: the model never saw your company's documents.

3. Fuzzy: it remembers patterns, not exact facts, so it can state wrong numbers confidently.

4. Uncitable: it can't point to where a fact came from.

External knowledge is information handed to the model at question time from your documents, databases, or APIs. It's current, private, exact, and traceable.

Flowchart

The model's job changes from "remember the answer" to "read these passages and answer from them." Reading is far more reliable than remembering.

Use cases

SituationUse thisWhy
"Explain what a refund is"The model's own knowledgeGeneral concept; no company facts needed
"What is ShopSphere's refund window?"External knowledge (RAG)A private fact the model never saw
"Write a polite apology email"The model's own knowledgeA writing skill, not a fact lookup
A bank bot asked for today's deposit rateExternal knowledge (the current rate sheet)Memory might return a rate from years ago

1.2 Three Things a Model Can't Know

Intuition: An old newspaper (printed before the news happened), a locked diary (never public), and a stock ticker (changes too fast).

ProblemMeaningShopSphere exampleRight fix
Knowledge cutoffTraining data ends at a dateReturn policy changed last monthRAG
Private dataNever on the public internetInternal exception rulesRAG
Fast-changing dataChanges within minutes or hoursPrices, stock, order statusLive lookup, not RAG

Key insight: a RAG index is a snapshot. The gap between "data changed" and "index updated" is the staleness window. If data changes faster than you can re-index, RAG will give confident but outdated answers.

Use cases

Data changes…Use thisWhy
Yearly or monthly (HR handbook)RAG with scheduled re-indexingCheap and stable
Daily or weekly (engineering wiki)RAG with incremental updatesRe-embed only what changed
Hourly (ShopSphere prices)Live API callAn index can't keep up
Every second (flight status, order tracking)Live lookup, alwaysMust be exact and current

Watch out: store an indexed_at timestamp on every chunk from day one. It costs nothing and lets you detect stale answers later.

1.3 Why "Just Put It in the Prompt" Has a Ceiling

Intuition: Handing someone a 1,000-page binder for every question. Fine at 10 pages; slow, costly, and error-prone at 1,000.

Prompt stuffing means pasting all your documents into the prompt. It's a great way to start: no infrastructure and quick results. But it hits limits:

LimitWhat happens
Context windowEventually your data doesn't fit
CostYou pay for the whole corpus on every question
SpeedMore input means slower responses
AttentionModels tend to use facts at the start and end of long inputs better than facts buried in the middle (the 2023 "Lost in the Middle" study by Liu et al.)
PermissionsEvery user's question sees every document

Simple cost math: cost per question ≈ input tokens × price. With stuffing, input tokens = your whole corpus. With RAG, input tokens = a few relevant chunks.

Use cases

SituationUse thisWhy
"Summarize this 20-page contract"Prompt stuffingThe task needs the whole document, and it fits
ShopSphere's 17,000 pagesRAGFar too big; each question needs a tiny slice
A demo for your manager next weekPrompt stuffingFastest way to prove the idea
A startup whose 30-page FAQ grew to 3,000 pagesSwitch from stuffing to RAGIt stopped fitting, and costs grew with every question

1.4 Grounding and Attribution: Requirements, Not Features

Intuition: A good journalist only reports what sources say (grounding) and names the source (attribution).

• Grounding: every claim in the answer is supported by the retrieved text.

• Attribution: the answer shows which passage supports each claim.

Why they're requirements:

• Trust: users can check the answer themselves.

• Debugging: you can tell whether the source was wrong or the model was.

• Compliance: legal, medical, and finance products often must show sources.

• Safety: a system allowed to say "I don't know" beats one that always guesses.

How it works (you'll build this in Topic 6):

1. Give every chunk a stable ID, like returns-v7#2.

2. Show those IDs to the model with the text.

3. Tell the model to cite IDs and to say "I don't know" when evidence is missing.

4. Check in code that every cited ID was actually retrieved. Models can invent plausible IDs.

Use cases

SituationUse thisWhy
Medical dosage lookupStrict grounding + required citations + refusal when unsureErrors can cause harm
Law firm research toolClause-level citations (e.g., MSA-2024 §7.2)Lawyers must verify in seconds
ShopSphere return policyGrounding + link to the help articleCustomers can confirm the rule
Brainstorming assistantLight grounding, citations optionalLow risk; creativity matters more