Topic 6: Hands-On Lab: ShopSphere's First Grounded Assistant
What you'll build: a small assistant that
1. answers policy questions with RAG and verified citations (Topics 1 and 2),
2. answers order questions with a live database lookup (Topic 3),
3. routes each question to the right path, or to both (Topic 3.5),
4. logs every step so failures can be diagnosed (Topic 4.1), and
5. is measured with a proper test set (Topic 4.4).
All code goes in module01/lab.py and uses the helpers from Topic 5. Run it from the rag-course/ folder with python -m module01.lab.
Step 1: Data
# module01/lab.py
import json
import re
import sqlite3
from qdrant_client import QdrantClient, models
from sentence_transformers import CrossEncoder
from common.embeddings import Embedder
from common.llm import ask, parse_json
# Help-center chunks (unstructured → RAG). Each has a stable ID for citations.
HELP_CHUNKS = [
{"id": "returns-v7#1", "text": "Unopened electronics can be returned within 15 days of delivery for a full refund."},
{"id": "returns-v7#2", "text": "Opened electronics can be returned within 7 days of delivery for store credit only."},
{"id": "returns-v7#3", "text": "Kitchen appliances, including blenders, can be returned within 30 days of delivery if unused."},
{"id": "shipping-v3#1", "text": "Standard shipping takes 3-5 business days. Express shipping takes 1-2 business days."},
{"id": "shipping-v3#2", "text": "Return shipping is free for orders above $50."},
{"id": "warranty-v2#1", "text": "All blenders include a 2-year warranty covering motor defects, not blade wear."},
]
# Orders (structured → live lookup, never embedded)
db = sqlite3.connect(":memory:")
db.execute("CREATE TABLE orders (id TEXT PRIMARY KEY, item TEXT, status TEXT, delivered_on TEXT)")
db.executemany("INSERT INTO orders VALUES (?, ?, ?, ?)", [
("4512", "X200 Blender", "delivered", "2026-09-05"),
("4513", "Noise-cancelling headphones", "shipped", None),
])
def get_order(order_id: str) -> dict | None:
row = db.execute("SELECT id, item, status, delivered_on FROM orders WHERE id = ?",
(order_id,)).fetchone() # parameterized: safe from SQL injection
return dict(zip(["id", "item", "status", "delivered_on"], row)) if row else NoneStep 2: Indexing pipeline (offline)
embedder = Embedder("hf") # BAAI/bge-small-en-v1.5
qdrant = QdrantClient(":memory:")
COLLECTION = "help_center"
def build_index(chunks: list[dict]) -> None:
vectors = embedder.embed_documents([c["text"] for c in chunks])
qdrant.create_collection(
collection_name=COLLECTION,
vectors_config=models.VectorParams(size=vectors.shape[1], distance=models.Distance.COSINE),
)
qdrant.upsert(
collection_name=COLLECTION,
points=[models.PointStruct(id=i, vector=v.tolist(), payload=c)
for i, (c, v) in enumerate(zip(chunks, vectors))],
)
print(f"Indexed {len(chunks)} chunks") # check: matches the source count?Step 3: Query pipeline (retrieve → rerank)
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
def retrieve(question: str, k: int = 5) -> list[dict]:
hits = qdrant.query_points(COLLECTION, query=embedder.embed_query(question).tolist(), limit=k).points
return [{**h.payload, "score": h.score} for h in hits]
def rerank(question: str, candidates: list[dict], k: int = 3) -> list[dict]:
if not candidates:
return []
scores = reranker.predict([(question, c["text"]) for c in candidates])
ranked = sorted(zip(scores, candidates), key=lambda x: x[0], reverse=True)
return [{**c, "rerank_score": float(s)} for s, c in ranked[:k]]Step 4: Router
ROUTER_SYSTEM = """Classify the customer question. Respond with ONLY a JSON object:
{"needs_policy": true or false, "order_id": "digits" or null}
- needs_policy: true if the answer depends on store rules (returns, shipping, warranty).
- order_id: the order number if the customer mentions one, otherwise null."""
def route(question: str) -> dict:
result = parse_json(ask(question, system=ROUTER_SYSTEM, json_mode=True), default={})
order_id = result.get("order_id")
# Safety net: trust digits found in the question over the model's guess
found = re.search(r"\b\d{3,}\b", question)
return {
"needs_policy": bool(result.get("needs_policy", True)),
"order_id": found.group(0) if found else (str(order_id) if order_id else None),
}Step 5: Assemble, generate, and verify citations
ANSWER_SYSTEM = """You are ShopSphere's support assistant.
Use ONLY the evidence provided. Each piece of evidence has an ID in square brackets.
- After every sentence, cite the supporting ID(s), e.g. [returns-v7#2] or [order-4512].
- Today's date is given; use it for any date calculations.
- If the evidence doesn't answer the question, reply exactly: "I don't have that information."
"""
def verify_citations(answer: str, allowed_ids: set[str]) -> dict:
cited = set(re.findall(r"\[([\w\-#.]+)\]", answer))
return {"cited": sorted(cited), "invalid": sorted(cited - allowed_ids)}
def assistant(question: str, today: str = "2026-09-16") -> dict:
trace = {"question": question, "route": route(question)}
evidence = {}
# Path 1: structured lookup (tool)
if order_id := trace["route"]["order_id"]:
order = get_order(order_id)
evidence[f"order-{order_id}"] = json.dumps(order) if order else "Order not found."
# Path 2: RAG
if trace["route"]["needs_policy"]:
candidates = retrieve(question)
best = rerank(question, candidates)
trace["retrieved"] = [(c["id"], round(c["score"], 3)) for c in candidates]
trace["reranked"] = [(c["id"], round(c["rerank_score"], 2)) for c in best]
evidence.update({c["id"]: c["text"] for c in best})
trace["in_prompt"] = list(evidence)
context = "\n".join(f"[{eid}] {text}" for eid, text in evidence.items())
answer = ask(f"Today's date: {today}\n\nEvidence:\n{context}\n\nQuestion: {question}",
system=ANSWER_SYSTEM)
check = verify_citations(answer, set(evidence))
refused = "don't have that information" in answer.lower()
trace["answer"] = answer
trace["citations"] = check
trace["safe_to_show"] = refused or (bool(check["cited"]) and not check["invalid"])
return traceStep 6: Try it
if __name__ == "__main__":
build_index(HELP_CHUNKS)
questions = [
"Can I return headphones I already opened?", # RAG only
"Where is my order 4513?", # tool only
"Can I still return the blender from order 4512?", # RAG + tool
"Do you ship to Mars?", # should refuse
]
for q in questions:
result = assistant(q)
print("\n" + "=" * 70)
print(json.dumps(result, indent=2))What to look for:
• Question 3 should combine the 30-day blender rule [returns-v7#3] with the delivery date [order-4512] and compute that the window is still open (delivered Sept 5, today is Sept 16).
• Question 4 should refuse rather than invent a shipping policy.
• Read the retrieved, reranked, and in_prompt fields. This trace is how you'll diagnose failures (Topic 4.1).
• Change LLM_PROVIDER in .env to compare Groq, Gemini, and Ollama on the same questions.
Step 7: Measure retrieval with a test set
TEST_SET = [
{"q": "How long can I return an unopened laptop?", "gold": "returns-v7#1", "type": "easy"},
{"q": "can i send back earbuds i already used", "gold": "returns-v7#2", "type": "paraphrase"},
{"q": "whats the return window for a blendr", "gold": "returns-v7#3", "type": "typo"},
{"q": "how quick is the fastest delivery option", "gold": "shipping-v3#1", "type": "paraphrase"},
{"q": "Do I pay postage when sending an item back?", "gold": "shipping-v3#2", "type": "paraphrase"},
{"q": "Is a broken blender blade covered?", "gold": "warranty-v2#1", "type": "easy"},
# Grow this to 50+ questions, including unanswerable ones, as the course progresses.
]
def recall_at_k(k: int = 3) -> None:
results = {}
for item in TEST_SET:
found = [c["id"] for c in retrieve(item["q"], k=k)]
hit = item["gold"] in found
results.setdefault(item["type"], []).append(hit)
if not hit:
print(f"MISS {item['q']!r} → got {found}") # a missed top-k failure (Topic 4.2)
total = [h for hits in results.values() for h in hits]
print(f"\nrecall@{k}: {sum(total)}/{len(total)}")
for qtype, hits in results.items():
print(f" {qtype:10s} {sum(hits)}/{len(hits)}")Add recall_at_k(k=1) and recall_at_k(k=3) to the main block. Then switch the embedder to Embedder("hf", "sentence-transformers/all-MiniLM-L6-v2") and compare. That's your first measured RAG experiment.
Module 1 Wrap-Up
Project Milestone: ShopSphere Architecture Decision Record
Write a one-page decision record before Module 2:
1. List every ShopSphere data source: structured or not, size, update frequency, who may see it.
2. Choose an approach for each (RAG, tool calling, long context, agentic search) with a one-sentence reason.
3. Describe the router: which question types go where.
4. Set a latency target and a cost-per-question target.
5. Extend the lab's test set to 50 questions, including 10 with no answer in the docs.
Interview Questions
1. What's the difference between parametric and external knowledge, and why does it matter for companies?
2. Your model has a huge context window. Why might you still choose RAG?
3. When would you choose fine-tuning over RAG? When would you use both?
4. Why is vector search the wrong tool for "How many orders shipped late last month?"
5. Walk through the RAG pipeline and name a silent failure at each stage.
6. How does prompt caching change the RAG vs long-context decision?
7. Your bot says "I don't know," but the docs contain the answer. How do you debug it?
8. Why are wrong-but-fluent answers more costly than refusals?
9. How would you build an evaluation set that avoids the "ten test questions" trap?
10. Design an assistant that answers both policy questions and live order questions.
Other LLM and Embedding Providers
The course helpers make it easy to add more providers. Add these branches to ask() in common/llm.py.
OpenAI (pip install openai, needs OPENAI_API_KEY)
if provider == "openai":
from openai import OpenAI
resp = OpenAI().responses.create(
model="gpt-5-mini", # check OpenAI's docs for current
instructions=system,
input=prompt,
)
return resp.output_text
Anthropic Claude (pip install anthropic, needs ANTHROPIC_API_KEY)
if provider == "anthropic":
from anthropic import Anthropic
resp = Anthropic().messages.create(
model="claude-sonnet-5",
max_tokens=1000,
system=system,
messages=[{"role": "user", "content": prompt}],
)
return resp.content[0].text
Embeddings from OpenAI (Anthropic doesn't offer its own embedding model; its docs point to partners such as Voyage AI)
from openai import OpenAI
result = OpenAI().embeddings.create(model="text-embedding-3-small",
input=["Free returns within 15 days."])
vector = [result.data](https://result.data)
[0].embedding # 1536 dimensions
| Provider | Chat | Embeddings | Free option | Offline |
| Groq | Yes | No | Free tier | No |
| Google Gemini | Yes | Yes | Free tier | No |
| Ollama | Yes | Yes | Fully free | Yes |
| Hugging Face models | Yes | Yes | Fully free | Yes |
| OpenAI | Yes | Yes | Paid | No |
| Anthropic | Yes | Via partners | Paid | No |
Coming Up in Module 2
Loading and parsing data from PDFs, Word documents, text files, websites, APIs, CSV, and Excel, and fixing the extraction failures from Topic 4.2.