CourseRAG · Module -1 :Foundations (What RAG Is and When It Is the Wrong Answer) · part 3 of 82
Part 3 · Module -1 :Foundations (What RAG Is and When It Is the Wrong Answer)

Topic 3: The Decision That Comes First

4 min read·21 Sept 2026

The most important idea in this module: RAG is one tool among several. Choose the architecture before writing code. No amount of chunking tuning fixes the wrong architecture.

3.1 RAG vs Long Context

Intuition: Long context is reading the whole book for every question. RAG is using the index to jump to the right page.

FactorLong contextRAG
Cost per questionHighLow
Is the fact available to the model?Always (everything is included)Only if retrieval finds it
Facts buried in the middleCan be used less reliablyFew chunks, so less of an issue
Reasoning across a whole documentStrongWeaker
Corpus size limitThe context windowPractically unlimited
Per-user permissionsHardEasy (filter chunks)

Prompt caching lets providers reuse a repeated prompt prefix at a lower price and latency. If many users query the same fixed documents, long context plus caching can compete with RAG. It helps less when documents differ per user or traffic is sparse. Pricing varies by provider, so benchmark.

Use cases

SituationUse thisWhy
"Do sections 3 and 9 of this contract conflict?"Long contextNeeds the whole document at once
Search across 50,000 contractsRAGNo context window can hold them
30-page product guide, thousands of daily questionsLong context + caching (benchmark vs RAG)Shared fixed prefix makes caching effective
Each user's private notesRAG with permission filtersDifferent data per user

3.2 RAG vs Fine-Tuning

Intuition: RAG is giving an employee a reference manual. Fine-tuning is sending them to a training course. Courses teach behaviour, not next week's facts.

You want to changeExampleBest tool
Knowledge"Our return window is 15 days"RAG
BehaviourAlways empathetic; follows a triage procedurePrompting first, then fine-tuning
FormatAlways valid JSON; house writing stylePrompting / structured output first

Fine-tuning is a poor way to add facts: they're fuzzy, need retraining to update, and can't be cited. Try in this order: prompting → RAG → fine-tuning. You can also combine them: fine-tune for style, use RAG for facts.

Use cases

SituationUse thisWhy
Answer from 10,000 policy documentsRAGKnowledge problem; docs change; citations needed
Strict brand voice in every replyPrompting, then fine-tuning if neededBehaviour problem
Classify tickets into 40 categories at high volumeFine-tuning a smaller modelRepetitive behaviour; lowers cost
A company fine-tuned on its catalog, and the model invented specsSwitch facts to RAGFine-tuning doesn't store facts reliably

3.3 RAG vs Tool Calling: Structured vs Unstructured Data

Intuition: You don't search a library for your bank balance. You ask the bank.

Tool calling: the LLM asks your code to run a function (like get_order_status("4512")), your code runs it, and the result goes back to the model.

Data typeExamplesUse thisWhy
StructuredOrders, inventory, CRM, pricesTool calling / SQL / APINeeds exact filters, counts, sums, freshness
UnstructuredPolicies, manuals, emailsRAGMeaning-based search over free text
Semi-structuredProduct listings with descriptions, JSON logsHybrid: filter fields, then semantic searchHas both exact fields and text

Why vector search fails on structured questions: "How many orders over $100 shipped last week?" needs filtering and counting. Similarity search returns text that looks similar, not an exact number.

Use cases

SituationUse thisWhy
"Where is my order?"Tool calling → orders databasePer-user, real-time, exact
"Revenue by region last quarter"Text-to-SQL with a read-only, validated queryAggregation over tables
"What does the warranty cover?"RAGFree-text policy
"Quiet blenders under $100"SQL filter on price + semantic search on reviewsMix of exact fields and meaning
Sales reports in CSV/ExcelLoad into a database or dataframe, then queryTables need computation

Watch out: if an LLM writes SQL, give it a read-only user and an allow-list of tables.

3.4 RAG vs Agentic Search

Intuition: RAG is your own filing cabinet. Agentic search is a research assistant who goes out, searches, reads, decides what to search next, and reports back.

FactorRAGAgentic search
FreshnessAs of the last re-indexLive
Speed and costFast, predictableSlower, variable
Multi-step questionsLimitedStrong
Must you store the data?YesNo
Main risksStale dataRunaway loops; malicious instructions hidden in web pages

Use cases

SituationUse thisWhy
ShopSphere help-center Q&ARAGOwned, stable, high volume
"What are competitors charging today?"Agentic web searchExternal, live, multi-step
Fast-changing codebaseAgentic search (search files, open, follow imports)Any index goes stale within hours
Ticketing system with its own search APIAgentic search via that APIData changes constantly; search already exists

Watch out: always cap the number of agent steps, and treat text from external pages as data, never as instructions.

3.5 Hybrid Strategies and the Three Deciding Variables

Intuition: A hospital triage desk routes each patient to the right specialist. Production assistants route each question the same way.

Three variables decide the architecture for each data source:

VariableLowHigh
Corpus sizeFits in the context windowFar beyond it
Update frequencyMonthlyEvery second
Query diversityA few repeated questionsThousands of different questions

Tip: if 80% of questions are the same 20, human-reviewed answers for those plus RAG for the rest is often cheaper and more accurate. A university chatbot did exactly this for fees, deadlines, and hostels.

ShopSphere, decided:

Data sourceDecisionWhy
Help-center articlesRAGText, large, weekly updates
Product manualsRAG (with table-aware parsing)Text-heavy PDFs, monthly updates
Orders databaseTool callingStructured, per-user, real-time
Pricing APITool callingChanges hourly
Competitor info (staff only)Agentic web searchExternal and live

"Can I return my blender from order 4512?" needs both the return policy (RAG) and the delivery date (tool call). A router sends it down both paths, and the LLM combines the results. You'll build exactly this in Topic 6.

Flowchart