CourseRAG · Module -1 :Foundations (What RAG Is and When It Is the Wrong Answer) · part 6 of 82
Part 6 · Module -1 :Foundations (What RAG Is and When It Is the Wrong Answer)

Topic 6: Hands-On Lab: ShopSphere's First Grounded Assistant

7 min read·21 Sept 2026

What you'll build: a small assistant that

1. answers policy questions with RAG and verified citations (Topics 1 and 2),

2. answers order questions with a live database lookup (Topic 3),

3. routes each question to the right path, or to both (Topic 3.5),

4. logs every step so failures can be diagnosed (Topic 4.1), and

5. is measured with a proper test set (Topic 4.4).

All code goes in module01/lab.py and uses the helpers from Topic 5. Run it from the rag-course/ folder with python -m module01.lab.

Step 1: Data

python
# module01/lab.py
import json
import re
import sqlite3

from qdrant_client import QdrantClient, models
from sentence_transformers import CrossEncoder

from common.embeddings import Embedder
from common.llm import ask, parse_json

# Help-center chunks (unstructured → RAG). Each has a stable ID for citations.
HELP_CHUNKS = [
    {"id": "returns-v7#1", "text": "Unopened electronics can be returned within 15 days of delivery for a full refund."},
    {"id": "returns-v7#2", "text": "Opened electronics can be returned within 7 days of delivery for store credit only."},
    {"id": "returns-v7#3", "text": "Kitchen appliances, including blenders, can be returned within 30 days of delivery if unused."},
    {"id": "shipping-v3#1", "text": "Standard shipping takes 3-5 business days. Express shipping takes 1-2 business days."},
    {"id": "shipping-v3#2", "text": "Return shipping is free for orders above $50."},
    {"id": "warranty-v2#1", "text": "All blenders include a 2-year warranty covering motor defects, not blade wear."},
]

# Orders (structured → live lookup, never embedded)
db = sqlite3.connect(":memory:")
db.execute("CREATE TABLE orders (id TEXT PRIMARY KEY, item TEXT, status TEXT, delivered_on TEXT)")
db.executemany("INSERT INTO orders VALUES (?, ?, ?, ?)", [
    ("4512", "X200 Blender", "delivered", "2026-09-05"),
    ("4513", "Noise-cancelling headphones", "shipped", None),
])

def get_order(order_id: str) -> dict | None:
    row = db.execute("SELECT id, item, status, delivered_on FROM orders WHERE id = ?",
                     (order_id,)).fetchone()          # parameterized: safe from SQL injection
    return dict(zip(["id", "item", "status", "delivered_on"], row)) if row else None

Step 2: Indexing pipeline (offline)

python
embedder = Embedder("hf")                              # BAAI/bge-small-en-v1.5
qdrant = QdrantClient(":memory:")
COLLECTION = "help_center"

def build_index(chunks: list[dict]) -> None:
    vectors = embedder.embed_documents([c["text"] for c in chunks])
    qdrant.create_collection(
        collection_name=COLLECTION,
        vectors_config=models.VectorParams(size=vectors.shape[1], distance=models.Distance.COSINE),
    )
    qdrant.upsert(
        collection_name=COLLECTION,
        points=[models.PointStruct(id=i, vector=v.tolist(), payload=c)
                for i, (c, v) in enumerate(zip(chunks, vectors))],
    )
    print(f"Indexed {len(chunks)} chunks")             # check: matches the source count?

Step 3: Query pipeline (retrieve → rerank)

python
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")

def retrieve(question: str, k: int = 5) -> list[dict]:
    hits = qdrant.query_points(COLLECTION, query=embedder.embed_query(question).tolist(), limit=k).points
    return [{**h.payload, "score": h.score} for h in hits]

def rerank(question: str, candidates: list[dict], k: int = 3) -> list[dict]:
    if not candidates:
        return []
    scores = reranker.predict([(question, c["text"]) for c in candidates])
    ranked = sorted(zip(scores, candidates), key=lambda x: x[0], reverse=True)
    return [{**c, "rerank_score": float(s)} for s, c in ranked[:k]]

Step 4: Router

python
ROUTER_SYSTEM = """Classify the customer question. Respond with ONLY a JSON object:
{"needs_policy": true or false, "order_id": "digits" or null}
- needs_policy: true if the answer depends on store rules (returns, shipping, warranty).
- order_id: the order number if the customer mentions one, otherwise null."""

def route(question: str) -> dict:
    result = parse_json(ask(question, system=ROUTER_SYSTEM, json_mode=True), default={})
    order_id = result.get("order_id")
    # Safety net: trust digits found in the question over the model's guess
    found = re.search(r"\b\d{3,}\b", question)
    return {
        "needs_policy": bool(result.get("needs_policy", True)),
        "order_id": found.group(0) if found else (str(order_id) if order_id else None),
    }

Step 5: Assemble, generate, and verify citations

python
ANSWER_SYSTEM = """You are ShopSphere's support assistant.
Use ONLY the evidence provided. Each piece of evidence has an ID in square brackets.
- After every sentence, cite the supporting ID(s), e.g. [returns-v7#2] or [order-4512].
- Today's date is given; use it for any date calculations.
- If the evidence doesn't answer the question, reply exactly: "I don't have that information."
"""

def verify_citations(answer: str, allowed_ids: set[str]) -> dict:
    cited = set(re.findall(r"\[([\w\-#.]+)\]", answer))
    return {"cited": sorted(cited), "invalid": sorted(cited - allowed_ids)}

def assistant(question: str, today: str = "2026-09-16") -> dict:
    trace = {"question": question, "route": route(question)}
    evidence = {}

    # Path 1: structured lookup (tool)
    if order_id := trace["route"]["order_id"]:
        order = get_order(order_id)
        evidence[f"order-{order_id}"] = json.dumps(order) if order else "Order not found."

    # Path 2: RAG
    if trace["route"]["needs_policy"]:
        candidates = retrieve(question)
        best = rerank(question, candidates)
        trace["retrieved"] = [(c["id"], round(c["score"], 3)) for c in candidates]
        trace["reranked"] = [(c["id"], round(c["rerank_score"], 2)) for c in best]
        evidence.update({c["id"]: c["text"] for c in best})

    trace["in_prompt"] = list(evidence)
    context = "\n".join(f"[{eid}] {text}" for eid, text in evidence.items())
    answer = ask(f"Today's date: {today}\n\nEvidence:\n{context}\n\nQuestion: {question}",
                 system=ANSWER_SYSTEM)

    check = verify_citations(answer, set(evidence))
    refused = "don't have that information" in answer.lower()
    trace["answer"] = answer
    trace["citations"] = check
    trace["safe_to_show"] = refused or (bool(check["cited"]) and not check["invalid"])
    return trace

Step 6: Try it

python
if __name__ == "__main__":
    build_index(HELP_CHUNKS)

    questions = [
        "Can I return headphones I already opened?",           # RAG only
        "Where is my order 4513?",                              # tool only
        "Can I still return the blender from order 4512?",     # RAG + tool
        "Do you ship to Mars?",                                 # should refuse
    ]
    for q in questions:
        result = assistant(q)
        print("\n" + "=" * 70)
        print(json.dumps(result, indent=2))

What to look for:

• Question 3 should combine the 30-day blender rule [returns-v7#3] with the delivery date [order-4512] and compute that the window is still open (delivered Sept 5, today is Sept 16).

• Question 4 should refuse rather than invent a shipping policy.

• Read the retrieved, reranked, and in_prompt fields. This trace is how you'll diagnose failures (Topic 4.1).

• Change LLM_PROVIDER in .env to compare Groq, Gemini, and Ollama on the same questions.

Step 7: Measure retrieval with a test set

python
TEST_SET = [
    {"q": "How long can I return an unopened laptop?",            "gold": "returns-v7#1", "type": "easy"},
    {"q": "can i send back earbuds i already used",               "gold": "returns-v7#2", "type": "paraphrase"},
    {"q": "whats the return window for a blendr",                 "gold": "returns-v7#3", "type": "typo"},
    {"q": "how quick is the fastest delivery option",             "gold": "shipping-v3#1", "type": "paraphrase"},
    {"q": "Do I pay postage when sending an item back?",          "gold": "shipping-v3#2", "type": "paraphrase"},
    {"q": "Is a broken blender blade covered?",                   "gold": "warranty-v2#1", "type": "easy"},
    # Grow this to 50+ questions, including unanswerable ones, as the course progresses.
]

def recall_at_k(k: int = 3) -> None:
    results = {}
    for item in TEST_SET:
        found = [c["id"] for c in retrieve(item["q"], k=k)]
        hit = item["gold"] in found
        results.setdefault(item["type"], []).append(hit)
        if not hit:
            print(f"MISS  {item['q']!r} → got {found}")   # a missed top-k failure (Topic 4.2)
    total = [h for hits in results.values() for h in hits]
    print(f"\nrecall@{k}: {sum(total)}/{len(total)}")
    for qtype, hits in results.items():
        print(f"  {qtype:10s} {sum(hits)}/{len(hits)}")

Add recall_at_k(k=1) and recall_at_k(k=3) to the main block. Then switch the embedder to Embedder("hf", "sentence-transformers/all-MiniLM-L6-v2") and compare. That's your first measured RAG experiment.

Module 1 Wrap-Up

Project Milestone: ShopSphere Architecture Decision Record

Write a one-page decision record before Module 2:

1. List every ShopSphere data source: structured or not, size, update frequency, who may see it.

2. Choose an approach for each (RAG, tool calling, long context, agentic search) with a one-sentence reason.

3. Describe the router: which question types go where.

4. Set a latency target and a cost-per-question target.

5. Extend the lab's test set to 50 questions, including 10 with no answer in the docs.

Interview Questions

1. What's the difference between parametric and external knowledge, and why does it matter for companies?

2. Your model has a huge context window. Why might you still choose RAG?

3. When would you choose fine-tuning over RAG? When would you use both?

4. Why is vector search the wrong tool for "How many orders shipped late last month?"

5. Walk through the RAG pipeline and name a silent failure at each stage.

6. How does prompt caching change the RAG vs long-context decision?

7. Your bot says "I don't know," but the docs contain the answer. How do you debug it?

8. Why are wrong-but-fluent answers more costly than refusals?

9. How would you build an evaluation set that avoids the "ten test questions" trap?

10. Design an assistant that answers both policy questions and live order questions.

Other LLM and Embedding Providers

The course helpers make it easy to add more providers. Add these branches to ask() in common/llm.py.

OpenAI (pip install openai, needs OPENAI_API_KEY)

python
if provider == "openai":  
  from openai import OpenAI  
  resp = OpenAI().responses.create(        
model="gpt-5-mini",                      # check OpenAI's docs for current
  instructions=system,    
   input=prompt,   
 )  
  return resp.output_text

Anthropic Claude (pip install anthropic, needs ANTHROPIC_API_KEY)

python
if provider == "anthropic": 
   from anthropic import Anthropic  
  resp = Anthropic().messages.create(  
      model="claude-sonnet-5",      
  max_tokens=1000,       
 system=system,      
  messages=[{"role": "user", "content": prompt}], 
   )   
 return resp.content[0].text

Embeddings from OpenAI (Anthropic doesn't offer its own embedding model; its docs point to partners such as Voyage AI)

python
from openai import OpenAI
result = OpenAI().embeddings.create(model="text-embedding-3-small",  
                              input=["Free returns within 15 days."])
vector = [result.data](https://result.data)

[0].embedding            # 1536 dimensions
ProviderChatEmbeddingsFree optionOffline
GroqYesNoFree tierNo
Google GeminiYesYesFree tierNo
OllamaYesYesFully freeYes
Hugging Face modelsYesYesFully freeYes
OpenAIYesYesPaidNo
AnthropicYesVia partnersPaidNo

Coming Up in Module 2

Loading and parsing data from PDFs, Word documents, text files, websites, APIs, CSV, and Excel, and fixing the extraction failures from Topic 4.2.