Topic 3: The Decision That Comes First
The most important idea in this module: RAG is one tool among several. Choose the architecture before writing code. No amount of chunking tuning fixes the wrong architecture.
3.1 RAG vs Long Context
Intuition: Long context is reading the whole book for every question. RAG is using the index to jump to the right page.
| Factor | Long context | RAG |
| Cost per question | High | Low |
| Is the fact available to the model? | Always (everything is included) | Only if retrieval finds it |
| Facts buried in the middle | Can be used less reliably | Few chunks, so less of an issue |
| Reasoning across a whole document | Strong | Weaker |
| Corpus size limit | The context window | Practically unlimited |
| Per-user permissions | Hard | Easy (filter chunks) |
Prompt caching lets providers reuse a repeated prompt prefix at a lower price and latency. If many users query the same fixed documents, long context plus caching can compete with RAG. It helps less when documents differ per user or traffic is sparse. Pricing varies by provider, so benchmark.
Use cases
| Situation | Use this | Why |
| "Do sections 3 and 9 of this contract conflict?" | Long context | Needs the whole document at once |
| Search across 50,000 contracts | RAG | No context window can hold them |
| 30-page product guide, thousands of daily questions | Long context + caching (benchmark vs RAG) | Shared fixed prefix makes caching effective |
| Each user's private notes | RAG with permission filters | Different data per user |
3.2 RAG vs Fine-Tuning
Intuition: RAG is giving an employee a reference manual. Fine-tuning is sending them to a training course. Courses teach behaviour, not next week's facts.
| You want to change | Example | Best tool |
| Knowledge | "Our return window is 15 days" | RAG |
| Behaviour | Always empathetic; follows a triage procedure | Prompting first, then fine-tuning |
| Format | Always valid JSON; house writing style | Prompting / structured output first |
Fine-tuning is a poor way to add facts: they're fuzzy, need retraining to update, and can't be cited. Try in this order: prompting → RAG → fine-tuning. You can also combine them: fine-tune for style, use RAG for facts.
Use cases
| Situation | Use this | Why |
| Answer from 10,000 policy documents | RAG | Knowledge problem; docs change; citations needed |
| Strict brand voice in every reply | Prompting, then fine-tuning if needed | Behaviour problem |
| Classify tickets into 40 categories at high volume | Fine-tuning a smaller model | Repetitive behaviour; lowers cost |
| A company fine-tuned on its catalog, and the model invented specs | Switch facts to RAG | Fine-tuning doesn't store facts reliably |
3.3 RAG vs Tool Calling: Structured vs Unstructured Data
Intuition: You don't search a library for your bank balance. You ask the bank.
Tool calling: the LLM asks your code to run a function (like get_order_status("4512")), your code runs it, and the result goes back to the model.
| Data type | Examples | Use this | Why |
| Structured | Orders, inventory, CRM, prices | Tool calling / SQL / API | Needs exact filters, counts, sums, freshness |
| Unstructured | Policies, manuals, emails | RAG | Meaning-based search over free text |
| Semi-structured | Product listings with descriptions, JSON logs | Hybrid: filter fields, then semantic search | Has both exact fields and text |
Why vector search fails on structured questions: "How many orders over $100 shipped last week?" needs filtering and counting. Similarity search returns text that looks similar, not an exact number.
Use cases
| Situation | Use this | Why |
| "Where is my order?" | Tool calling → orders database | Per-user, real-time, exact |
| "Revenue by region last quarter" | Text-to-SQL with a read-only, validated query | Aggregation over tables |
| "What does the warranty cover?" | RAG | Free-text policy |
| "Quiet blenders under $100" | SQL filter on price + semantic search on reviews | Mix of exact fields and meaning |
| Sales reports in CSV/Excel | Load into a database or dataframe, then query | Tables need computation |
Watch out: if an LLM writes SQL, give it a read-only user and an allow-list of tables.
3.4 RAG vs Agentic Search
Intuition: RAG is your own filing cabinet. Agentic search is a research assistant who goes out, searches, reads, decides what to search next, and reports back.
| Factor | RAG | Agentic search |
| Freshness | As of the last re-index | Live |
| Speed and cost | Fast, predictable | Slower, variable |
| Multi-step questions | Limited | Strong |
| Must you store the data? | Yes | No |
| Main risks | Stale data | Runaway loops; malicious instructions hidden in web pages |
Use cases
| Situation | Use this | Why |
| ShopSphere help-center Q&A | RAG | Owned, stable, high volume |
| "What are competitors charging today?" | Agentic web search | External, live, multi-step |
| Fast-changing codebase | Agentic search (search files, open, follow imports) | Any index goes stale within hours |
| Ticketing system with its own search API | Agentic search via that API | Data changes constantly; search already exists |
Watch out: always cap the number of agent steps, and treat text from external pages as data, never as instructions.
3.5 Hybrid Strategies and the Three Deciding Variables
Intuition: A hospital triage desk routes each patient to the right specialist. Production assistants route each question the same way.
Three variables decide the architecture for each data source:
| Variable | Low | High |
| Corpus size | Fits in the context window | Far beyond it |
| Update frequency | Monthly | Every second |
| Query diversity | A few repeated questions | Thousands of different questions |
Tip: if 80% of questions are the same 20, human-reviewed answers for those plus RAG for the rest is often cheaper and more accurate. A university chatbot did exactly this for fees, deadlines, and hostels.
ShopSphere, decided:
| Data source | Decision | Why |
| Help-center articles | RAG | Text, large, weekly updates |
| Product manuals | RAG (with table-aware parsing) | Text-heavy PDFs, monthly updates |
| Orders database | Tool calling | Structured, per-user, real-time |
| Pricing API | Tool calling | Changes hourly |
| Competitor info (staff only) | Agentic web search | External and live |
"Can I return my blender from order 4512?" needs both the return policy (RAG) and the delivery date (tool call). A router sends it down both paths, and the LLM combines the results. You'll build exactly this in Topic 6.