Topic 1: Why RAG Exists (The Core Problem)
Module 1: Foundations (What RAG Is and When It Is the Wrong Answer)
By the end of this module, you'll be able to:
- Explain what problem RAG solves, and what it doesn't
- Draw the RAG pipeline from memory and say where each stage breaks
- Choose between RAG, long context, fine-tuning, tool calling, and agentic search, with reasons
- Name a failure precisely so you can fix it
- Set up the tech stack used for the rest of the course and build your first grounded RAG assistant
How this module is organized
| Topic | What you'll learn | Style |
| 1. Why RAG Exists | The problem RAG solves | Concepts |
| 2. Anatomy of a RAG System | The pipeline and where it breaks | Concepts |
| 3. The Decision That Comes First | When to use RAG, and when not to | Concepts |
| 4. Failure Modes, Named Early | How to diagnose bad answers | Concepts |
| 5. The Course Tech Stack | LLMs, embedding models, rerankers, vector DBs, frameworks | Setup + examples |
| 6. Hands-On Lab | Build ShopSphere's first grounded assistant | Code |
Topics 1-4 are about thinking, so there's no code. You'll apply every idea in Topic 6.
The Running Project: ShopSphere Support Assistant
Throughout the course you'll build one assistant for ShopSphere, a fictional online store. It has four data sources:
| Data source | Type | Size | Changes |
| Help-center articles (returns, shipping) | Text | ~2,000 pages | Weekly |
| Product manuals | PDFs | ~15,000 pages | Monthly |
| Orders database | Table | Millions of rows | Every second |
| Pricing API | API | Thousands of products | Hourly |
Keep these in mind. By the end of Topic 3, you'll know exactly which approach each one needs.
Topic 1: Why RAG Exists (The Core Problem)
1.1 Parametric Knowledge vs External Knowledge
Intuition: An LLM answering from memory is a student taking a closed-book exam. RAG turns it into an open-book exam: the student still needs to be smart, but now they can look things up.
Parametric knowledge is what the model learned in training, stored in its weights. It has four limits:
1. Frozen: it stops at the training date.
2. Public only: the model never saw your company's documents.
3. Fuzzy: it remembers patterns, not exact facts, so it can state wrong numbers confidently.
4. Uncitable: it can't point to where a fact came from.
External knowledge is information handed to the model at question time from your documents, databases, or APIs. It's current, private, exact, and traceable.
The model's job changes from "remember the answer" to "read these passages and answer from them." Reading is far more reliable than remembering.
Use cases
| Situation | Use this | Why |
| "Explain what a refund is" | The model's own knowledge | General concept; no company facts needed |
| "What is ShopSphere's refund window?" | External knowledge (RAG) | A private fact the model never saw |
| "Write a polite apology email" | The model's own knowledge | A writing skill, not a fact lookup |
| A bank bot asked for today's deposit rate | External knowledge (the current rate sheet) | Memory might return a rate from years ago |
1.2 Three Things a Model Can't Know
Intuition: An old newspaper (printed before the news happened), a locked diary (never public), and a stock ticker (changes too fast).
| Problem | Meaning | ShopSphere example | Right fix |
| Knowledge cutoff | Training data ends at a date | Return policy changed last month | RAG |
| Private data | Never on the public internet | Internal exception rules | RAG |
| Fast-changing data | Changes within minutes or hours | Prices, stock, order status | Live lookup, not RAG |
Key insight: a RAG index is a snapshot. The gap between "data changed" and "index updated" is the staleness window. If data changes faster than you can re-index, RAG will give confident but outdated answers.
Use cases
| Data changes… | Use this | Why |
| Yearly or monthly (HR handbook) | RAG with scheduled re-indexing | Cheap and stable |
| Daily or weekly (engineering wiki) | RAG with incremental updates | Re-embed only what changed |
| Hourly (ShopSphere prices) | Live API call | An index can't keep up |
| Every second (flight status, order tracking) | Live lookup, always | Must be exact and current |
Watch out: store an indexed_at timestamp on every chunk from day one. It costs nothing and lets you detect stale answers later.
1.3 Why "Just Put It in the Prompt" Has a Ceiling
Intuition: Handing someone a 1,000-page binder for every question. Fine at 10 pages; slow, costly, and error-prone at 1,000.
Prompt stuffing means pasting all your documents into the prompt. It's a great way to start: no infrastructure and quick results. But it hits limits:
| Limit | What happens |
| Context window | Eventually your data doesn't fit |
| Cost | You pay for the whole corpus on every question |
| Speed | More input means slower responses |
| Attention | Models tend to use facts at the start and end of long inputs better than facts buried in the middle (the 2023 "Lost in the Middle" study by Liu et al.) |
| Permissions | Every user's question sees every document |
Simple cost math: cost per question ≈ input tokens × price. With stuffing, input tokens = your whole corpus. With RAG, input tokens = a few relevant chunks.
Use cases
| Situation | Use this | Why |
| "Summarize this 20-page contract" | Prompt stuffing | The task needs the whole document, and it fits |
| ShopSphere's 17,000 pages | RAG | Far too big; each question needs a tiny slice |
| A demo for your manager next week | Prompt stuffing | Fastest way to prove the idea |
| A startup whose 30-page FAQ grew to 3,000 pages | Switch from stuffing to RAG | It stopped fitting, and costs grew with every question |
1.4 Grounding and Attribution: Requirements, Not Features
Intuition: A good journalist only reports what sources say (grounding) and names the source (attribution).
• Grounding: every claim in the answer is supported by the retrieved text.
• Attribution: the answer shows which passage supports each claim.
Why they're requirements:
• Trust: users can check the answer themselves.
• Debugging: you can tell whether the source was wrong or the model was.
• Compliance: legal, medical, and finance products often must show sources.
• Safety: a system allowed to say "I don't know" beats one that always guesses.
How it works (you'll build this in Topic 6):
1. Give every chunk a stable ID, like returns-v7#2.
2. Show those IDs to the model with the text.
3. Tell the model to cite IDs and to say "I don't know" when evidence is missing.
4. Check in code that every cited ID was actually retrieved. Models can invent plausible IDs.
Use cases
| Situation | Use this | Why |
| Medical dosage lookup | Strict grounding + required citations + refusal when unsure | Errors can cause harm |
| Law firm research tool | Clause-level citations (e.g., MSA-2024 §7.2) | Lawyers must verify in seconds |
| ShopSphere return policy | Grounding + link to the help article | Customers can confirm the rule |
| Brainstorming assistant | Light grounding, citations optional | Low risk; creativity matters more |