Topic 2: Anatomy of a RAG System
2.1 Two Pipelines: Indexing Time vs Query Time
Intuition: A library's back office catalogs books slowly and carefully overnight. The front desk answers visitors in seconds.
| Indexing pipeline (offline) | Query pipeline (online) | |
| Runs when | Documents are added or changed | A user asks a question |
| Speed needed | Minutes or hours is fine | Seconds or less |
| Optimize for | Quality and completeness | Speed and cost per question |
| Heavy work allowed? | Yes (OCR, slow parsers) | Only if fast |
Rule of thumb: do expensive work once at indexing time so every query stays fast.
Golden rule: use the same embedding model in both pipelines. Vectors from different models can't be compared.
Use cases
| Situation | Use this | Why |
| 15,000 pages of PDF manuals | Heavy parsing and OCR at indexing time | Slow is fine offline; quality pays off on every query |
| A user uploads a PDF and wants to chat now | A lighter, faster indexing path | The user is waiting |
| Help center updated weekly | Weekly scheduled re-index | Matches the update rhythm |
2.2 The Nine Stages
Intuition: A restaurant kitchen: prep ingredients, store them labeled, pick the right ones per order, choose the best, plate, cook, and attach the menu card.
INDEXING
[Parse] → [Chunk] → [Embed] → [(Index)]
│
┆
▼
QUERY
[Question] → [Retrieve Top-20] → [Rerank Top-5] → [Assemble Prompt] → [Generate] → [Cite & Verify]| # | Stage | What it does | Covered in |
| 1 | Parse | Extract clean text and tables from files | Module 2 |
| 2 | Chunk | Split text into small, searchable pieces | Module 3 |
| 3 | Embed | Turn each chunk into a vector (numbers that capture meaning) | Module 4 |
| 4 | Index | Store vectors so similar ones are found quickly | Module 5 |
| 5 | Retrieve | Find the chunks most similar to the question | Module 6 |
| 6 | Rerank | Re-score candidates with a slower, more accurate model | Module 7 |
| 7 | Assemble | Build the prompt: instructions + chunks + question | Module 8 |
| 8 | Generate | The LLM writes the answer | Module 8 |
| 9 | Cite | Link claims to sources and verify them | Module 8 |
Why retrieve and rerank? Retrieval is fast but rough; reranking is slow but sharp. Retrieve broadly (20-50 candidates), rerank narrowly (3-8). You get both speed and accuracy.
Use cases
| Situation | Use this | Why |
| 200 short FAQ entries | Skip chunking and reranking | Entries are already small; corpus is tiny |
| ShopSphere manuals | All nine stages | Many similar passages; ranking precision matters |
| Voice assistant | All stages, with a small reranker or none | Every millisecond counts |
| Legal research | All stages + strict citation checks | Precision and traceability are essential |
2.3 Where Each Stage Silently Fails
Intuition: A leaky pipe inside a wall: no alarm, just slowly worse results. The LLM smooths over bad input with fluent language, hiding the problem.
| Stage | Silent failure | What users see |
| Parse | Scanned PDFs return empty text; tables become jumbled | "No information" for docs that exist; wrong numbers |
| Chunk | A rule is separated from its exception | Half-right answers missing the "unless…" |
| Embed | Model doesn't understand your jargon | Relevant docs never found |
| Index | New docs never added; deleted docs remain | Outdated or contradictory answers |
| Retrieve | Right chunk ranks 15th when only 5 are kept | "I don't know" when the answer exists |
| Rerank | Reranker pushes the right chunk out | Plausible but wrong answers |
| Assemble | Key chunk cut off by the token limit | Model ignores the key fact |
| Generate | Model mixes context with its own memory | Confident wrong answer |
| Cite | Citation doesn't support the claim | Users trust something false |
Real story: an insurance company found that 30% of its policy PDFs were scanned images parsed as empty text. No error was ever raised. The bot had quietly answered from the other 70% for months.
Watch out: add a simple check after each stage: text length after parsing, document counts in source vs index, and low retrieval scores. Monitoring "the API returned 200" isn't enough.
2.4 Latency and Cost Budget
Intuition: A trip budget. Generation is usually the most expensive hotel on the trip.
| Stage | Runs at | Typical time (varies, measure yours) | Main cost |
| Parse, chunk, embed documents | Indexing | Minutes to hours (batch) | Compute, one-time per document |
| Embed query | Query | ~10-100 ms | Tiny |
| Retrieve | Query | ~5-100 ms | Vector DB hosting |
| Rerank | Query | ~50-500 ms | Reranker compute |
| Generate | Query | ~0.5-10+ s | LLM tokens (usually the biggest) |
Key levels: send fewer, better chunks; stream responses; use smaller models for simple questions; cache frequent questions; limit rerank candidates.
Real story: a bot took 6 seconds per answer. Reranking 25 candidates instead of 100 and sending 5 chunks instead of 15 brought it to about 2.5 seconds, with the same answer quality.
Use cases
| Situation | Use this | Why |
| 200 short FAQ entries | Skip chunking and reranking | Entries are already small; corpus is tiny |
| ShopSphere manuals | All nine stages | Many similar passages; ranking precision matters |
| Voice assistant | All stages, with a small reranker or none | Every millisecond counts |
| Legal research | All stages + strict citation checks | Precision and traceability are essential |