Overview
Large Language Models: Understand, Direct, Adapt, Evaluate, Operate
Course Overview
Most developers use large language models as a black box: write a prompt, check that the answer looks right, and ship it. This course teaches you what is actually happening inside, so you can make the model do what you want, measure whether it works, and run it safely and cheaply in production.
Across 14 modules and a capstone project, you build one system from start to finish: a support assistant for Brightlane, a fictional software company. The assistant triages tickets in five languages, answers questions from a help center, drafts replies for a person to review, and takes safe actions as an agent.
Every design choice comes with numbers you can check: token counts, cost, speed, and accuracy. You also work with TinyLM, a small model that runs on a laptop, so you can watch next-token prediction, sampling, and fine-tuning happen for real. All the examples work with Groq, Gemini, or Ollama, and each has a free option.
What You'll Learn
- How LLMs work: next-token prediction, tokens, context windows, sampling, and cost
- Prompting and reasoning techniques, and how to measure whether they help
- Reliable structured output, tool calling, context engineering, and agents
- When and how to fine-tune a model, including LoRA and DPO
- How to evaluate LLM systems, including using an LLM as a judge
- Security and safety: prompt injection, data leaks, and personal data handling
- Working with images, documents, audio, and video
- Running LLMs in production: cost, caching, fallbacks, monitoring, and upgrades
Prerequisites
- Working Python and basic use of the command line
- No machine learning background, GPU, or paid account needed
- Python 3.11 and about 3 GB of free disk space
- Optional: a free Groq or Gemini API key, or Ollama installed locally, to run real models
What you will learn
- 80 written parts, yours for good
- 27h 0m of reading, measured not estimated
- Written for the advanced level
- Every future revision included
Curriculum
15 sections · 80 parts · 27h 0m01Module 1: What a Large Language Model Actually Is7 parts · 101 min
- Part A: The Project and the ToolkitFree19 min
- Part B: Next-Token Prediction, the Only Thing the Model DoesFree14 min
- Part C: What Pretraining LearnedFree9 min
- Part D: Reading a Model SpecFree8 min
- Part E: From Base Model to Assistant: The LifecycleFree12 min
- Part F: Capabilities and Limits, MeasuredFree17 min
- Part G: The Model LandscapeFree22 min
02Module 2 : Tokens -context-cost4 parts · 85 min
- Part A: Tokenization in practiceFree30 min
- Part B: The context windowFree19 min
- Part C: Managing contextFree11 min
- Part D: Cost mechanicsFree25 min
03Module 3: Inference Behaviour and Decoding Control4 parts · 81 min
- Part A: How a response is producedFree24 min
- Part B: Sampling parametersFree22 min
- Part C: Constrained generationFree13 min
- Part D: Reasoning-model behaviourFree22 min
04Module 4: Prompt Engineering Fundamentals6 parts · 87 min
- Part A: Anatomy of a promptFree18 min
- Part B: Build the test set before you tuneFree18 min
- Part C: Core techniquesFree11 min
- Part D: Formatting and trust boundariesFree11 min
- Part E: Iteration disciplineFree10 min
- Part F: Anti-patternsFree19 min
05Module 5: Reasoning and Advanced Prompting5 parts · 107 min
- Part A: Eliciting reasoningFree26 min
- Part B: Sampling-based improvementFree24 min
- Part C: DecompositionFree19 min
- Part D: Self-correctionFree8 min
- Part E: Programmatic promptingFree30 min
06Module 6: Structured Outputs and Tool Use3 parts · 103 min
- Part A: Getting reliable structureFree45 min
- Part B: Tool and function callingFree34 min
- Part C: Model output as untrusted inputFree24 min
07Module 7: Context Engineering6 parts · 111 min
- Part A: The disciplineFree15 min
- Part B: Retrieved knowledgeFree15 min
- Part C: Assembling a contextFree25 min
- Part D: MemoryFree15 min
- Part E: Failure modesFree12 min
- Part F: OptimizationFree29 min
08Module 8: Agents8 parts · 104 min
- Part 1: Workflows and agentsFree6 min
- Part 2: The core loopFree24 min
- Part 3: Stopping, budgets, and recoveryFree8 min
- Part 4: Is the agent worth it?Free8 min
- Part 5: Reliability in the loopFree7 min
- Part 6: CapabilitiesFree14 min
- Part 7: Multi-agent systemsFree14 min
- Part 8: Evaluating trajectoriesFree23 min
09Module 9: Adaptation: Fine-Tuning and Customization7 parts · 123 min
- Part 1: The decision firstFree10 min
- Part 2: Measure before you trainFree22 min
- Part 3: DataFree12 min
- Part 4: MethodsFree25 min
- Part 5: Preference and reasoning trainingFree26 min
- Part 6: Serving many adapters over one base modelFree5 min
- Part 7: VerificationFree23 min
10Module 10: Evaluation7 parts · 146 min
- Part A: Why evaluation is the bottleneckFree11 min
- Part B: Constructing evalsFree36 min
- Part C: Noise, sample size, and comparing systemsFree13 min
- Part D: Human evaluationFree17 min
- Part E: LLM-as-judgeFree17 min
- Part F: System-level evaluationFree16 min
- Part G: Operating evalsFree36 min
11Module 11: Safety, Security, and Alignment in Practice4 parts · 101 min
- Part A: Model-level alignmentFree21 min
- Part B: The attack surfaceFree22 min
- Part C: Defenses, measuredFree27 min
- Part D: Responsible deploymentFree31 min
12Module 12: Multimodal Models4 parts · 116 min
- Part 1: Vision-language modelsFree42 min
- Part 2: DocumentsFree27 min
- Part 3: Other modalitiesFree25 min
- Part 4: Multimodal system designFree22 min
13Module 13: Deployment, Operations, and Economics5 parts · 131 min
- Part A: Access modelFree21 min
- Part B: Self-hostingFree30 min
- Part C: Operating in productionFree32 min
- Part D: ObservabilityFree19 min
- Part E: Change managementFree29 min
14Module 14: Applications, Patterns, and Frontiers4 parts · 107 min
- Part A: Proven Application PatternsFree39 min
- Part B: Design PrinciplesFree19 min
- Part C: JudgementFree19 min
- Part D: FrontierFree30 min
15Capstone Project: Four Tracks6 parts · 117 min
- Part A: Choosing a TrackFree8 min
- Part B: Shared RequirementsFree43 min
- Part C: Product TrackFree10 min
- Part D: Agent TrackFree15 min
- Part E: Adaptation TrackFree11 min
- Part F: Evaluation TrackFree30 min
By the end of this module, you'll have:
- Run a real language model (TinyLM, 1.07 million parameters) on your own CPU, read its next-token probabilities, and generated a support reply one token at a time.
- Read a real training log, pointed at the step where the model started to overfit, and measured how much of its output is memorized text.
Instructor
Naresh Edagotti
Course instructor
Reviews
to review this course once you have finished it.