CourseRAG · Module -1 :Foundations (What RAG Is and When It Is the Wrong Answer) · part 2 of 82
Part 2 · Module -1 :Foundations (What RAG Is and When It Is the Wrong Answer)

Topic 2: Anatomy of a RAG System

3 min read·21 Sept 2026

2.1 Two Pipelines: Indexing Time vs Query Time

Intuition: A library's back office catalogs books slowly and carefully overnight. The front desk answers visitors in seconds.

 Indexing pipeline (offline)Query pipeline (online)
Runs whenDocuments are added or changedA user asks a question
Speed neededMinutes or hours is fineSeconds or less
Optimize forQuality and completenessSpeed and cost per question
Heavy work allowed?Yes (OCR, slow parsers)Only if fast

Rule of thumb: do expensive work once at indexing time so every query stays fast.

Golden rule: use the same embedding model in both pipelines. Vectors from different models can't be compared.

Use cases

SituationUse thisWhy
15,000 pages of PDF manualsHeavy parsing and OCR at indexing timeSlow is fine offline; quality pays off on every query
A user uploads a PDF and wants to chat nowA lighter, faster indexing pathThe user is waiting
Help center updated weeklyWeekly scheduled re-indexMatches the update rhythm

2.2 The Nine Stages

Intuition: A restaurant kitchen: prep ingredients, store them labeled, pick the right ones per order, choose the best, plate, cook, and attach the menu card.

text
INDEXING

[Parse] → [Chunk] → [Embed] → [(Index)]
                                      │
                                      ┆
                                      ▼
QUERY

[Question] → [Retrieve Top-20] → [Rerank Top-5] → [Assemble Prompt] → [Generate] → [Cite & Verify]
#StageWhat it doesCovered in
1ParseExtract clean text and tables from filesModule 2
2ChunkSplit text into small, searchable piecesModule 3
3EmbedTurn each chunk into a vector (numbers that capture meaning)Module 4
4IndexStore vectors so similar ones are found quicklyModule 5
5RetrieveFind the chunks most similar to the questionModule 6
6RerankRe-score candidates with a slower, more accurate modelModule 7
7AssembleBuild the prompt: instructions + chunks + questionModule 8
8GenerateThe LLM writes the answerModule 8
9CiteLink claims to sources and verify themModule 8

Why retrieve and rerank? Retrieval is fast but rough; reranking is slow but sharp. Retrieve broadly (20-50 candidates), rerank narrowly (3-8). You get both speed and accuracy.

Use cases

SituationUse thisWhy
200 short FAQ entriesSkip chunking and rerankingEntries are already small; corpus is tiny
ShopSphere manualsAll nine stagesMany similar passages; ranking precision matters
Voice assistantAll stages, with a small reranker or noneEvery millisecond counts
Legal researchAll stages + strict citation checksPrecision and traceability are essential

2.3 Where Each Stage Silently Fails

Intuition: A leaky pipe inside a wall: no alarm, just slowly worse results. The LLM smooths over bad input with fluent language, hiding the problem.

StageSilent failureWhat users see
ParseScanned PDFs return empty text; tables become jumbled"No information" for docs that exist; wrong numbers
ChunkA rule is separated from its exceptionHalf-right answers missing the "unless…"
EmbedModel doesn't understand your jargonRelevant docs never found
IndexNew docs never added; deleted docs remainOutdated or contradictory answers
RetrieveRight chunk ranks 15th when only 5 are kept"I don't know" when the answer exists
RerankReranker pushes the right chunk outPlausible but wrong answers
AssembleKey chunk cut off by the token limitModel ignores the key fact
GenerateModel mixes context with its own memoryConfident wrong answer
CiteCitation doesn't support the claimUsers trust something false

Real story: an insurance company found that 30% of its policy PDFs were scanned images parsed as empty text. No error was ever raised. The bot had quietly answered from the other 70% for months.

Watch out: add a simple check after each stage: text length after parsing, document counts in source vs index, and low retrieval scores. Monitoring "the API returned 200" isn't enough.

2.4 Latency and Cost Budget

Intuition: A trip budget. Generation is usually the most expensive hotel on the trip.

StageRuns atTypical time (varies, measure yours)Main cost
Parse, chunk, embed documentsIndexingMinutes to hours (batch)Compute, one-time per document
Embed queryQuery~10-100 msTiny
RetrieveQuery~5-100 msVector DB hosting
RerankQuery~50-500 msReranker compute
GenerateQuery~0.5-10+ sLLM tokens (usually the biggest)

Key levels: send fewer, better chunks; stream responses; use smaller models for simple questions; cache frequent questions; limit rerank candidates.

Real story: a bot took 6 seconds per answer. Reranking 25 candidates instead of 100 and sending 5 chunks instead of 15 brought it to about 2.5 seconds, with the same answer quality.

Use cases

SituationUse thisWhy
200 short FAQ entriesSkip chunking and rerankingEntries are already small; corpus is tiny
ShopSphere manualsAll nine stagesMany similar passages; ranking precision matters
Voice assistantAll stages, with a small reranker or noneEvery millisecond counts
Legal researchAll stages + strict citation checksPrecision and traceability are essential