Overview
The classical MLOps engineer interview, answered. Fourteen volumes cover the production mindset, data engineering for MLOps, experiment tracking and the model lifecycle, testing ML systems, CI/CD and continuous training, model serving and inference, monitoring and reliability, model governance, ML security, cloud MLOps, performance and cost engineering, production ML system design, debugging and failure diagnosis, and company-style interview rounds. Every answer is written in the first person with a concrete example, and many carry a diagram.
What you will learn
- Explain why MLOps exists and what it adds to DevOps
- Build data pipelines and track experiments and models
- Test ML systems and automate CI/CD and continuous training
- Serve models and monitor them for drift and failures
- Govern and secure models in production
- Run MLOps on the major clouds and control cost
- Design production ML systems and debug them under pressure
- Handle company-style MLOps interview rounds
Included with the kit
- 14 written parts, yours for good
- 20h 36m of reading, measured not estimated
- Written for the intermediate level
- Every future revision included
- UPI, cards and netbanking
Prerequisites
- Software engineering or data engineering experience
- Basic machine learning concepts
Curriculum
14 sections · 14 parts · 20h 36m01Volume 1: Foundations & Production Mindset1 part · 82 min
ML lifecycle, CRISP-DM, DevOps vs DataOps vs MLOps vs LLMOps, ML platform evolution, maturity levels, orchestration (Airflow/Kubeflow/Dagster/Prefect), reproducibility, env mgmt (Docker/Conda/Poetry), Git/GitOps, Terraform/IaC, secrets mgmt, repo structure
- Foundations & Production Mindset82 min
02Volume 2: Data Engineering for MLOps1 part · 106 min
Ingestion (batch/streaming/Kafka/Spark), Delta Lake/Iceberg, DVC/LakeFS, data contracts, lineage, schema evolution, feature engineering, feature stores (Feast/Tecton), online vs offline, point-in-time correctness, data validation (Great Expectations/TFDV/Evidently), feature freshness, drift sources
- Data Engineering for MLOps106 min
03Volume 3: Experiment Tracking & Model Lifecycle1 part · 82 min
MLflow/W&B/Comet/Neptune, artifacts, lineage, metadata, registry, champion/challenger, approval, promotion, rollback, reproducibility, versioning, auditability
- Experiment Tracking & Model Lifecycle82 min
04Volume 4: Testing ML Systems1 part · 81 min
Unit/data/feature/pipeline/integration/regression/shadow/smoke/acceptance testing, golden datasets, model validation, CI testing
- Testing ML Systems81 min
05Volume 5: CI/CD/CT for ML1 part · 58 min
GitHub Actions/GitLab CI/Azure DevOps/Jenkins, continuous training/evaluation/deployment/delivery, model promotion, rollback automation
- CI/CD/CT for ML58 min
06Volume 6: Model Serving & Inference1 part · 68 min
FastAPI/gRPC/REST, Docker/K8s/Helm, KEDA/HPA, GPU scheduling, BentoML/Triton/Seldon/TorchServe, canary/blue-green/shadow, autoscaling, ONNX/TensorRT/OpenVINO, quantization, batching, rate limiting, caching, auth
- Model Serving & Inference68 min
07Volume 7: Monitoring & Reliability1 part · 85 min
Prometheus/Grafana/Datadog/OpenTelemetry, logging/tracing/metrics, latency/CPU/GPU/memory, business metrics, concept/prediction/feature drift, training-serving skew, alerting, SHAP/LIME, fairness/bias, RCA
- Monitoring & Reliability85 min
08Volume 8: Model Governance1 part · 38 min
Compliance, lineage, auditability, model cards, approval, RBAC, version control, access control, PII, data retention, documentation
- Model Governance38 min
09Volume 9: ML Security1 part · 83 min
Model extraction, data/model poisoning, adversarial examples, secrets, API abuse, dependency scanning, container security, supply chain. Prompt injection is compared briefly here and covered in full in the AI Security kit.
- ML Security83 min
10Volume 10: Cloud MLOps1 part · 87 min
AWS (SageMaker/S3/Lambda/EKS/ECS/ECR), Azure (Azure ML/AKS/Blob), GCP (Vertex AI/Cloud Run/BigQuery/GKE), multi-cloud, hybrid cloud
- Cloud MLOps87 min
11Volume 11: Performance & Cost Engineering1 part · 76 min
GPU/CPU optimization, autoscaling, compression/distillation/quantization, batching, caching, spot instances, scheduling, resource allocation, profiling, benchmarking
- Performance & Cost Engineering76 min
12Volume 12: Production ML System Design1 part · 104 min
Recsys, fraud detection, search, ads, ranking, OCR, medical AI, forecasting, vision, speech, real-time/streaming ML, multi-region, HA, disaster recovery
- Production ML System Design104 min
13Volume 13: Debugging & Failure Diagnosis1 part · 175 min
Broken pipelines, OOM, CrashLoopBackOff, feature drift, registry failures, deployment failures, timeouts, schema mismatch, version conflicts, training-serving mismatch, silent failures, RCA
- Debugging & Failure Diagnosis175 min
14Volume 14: Company-Style Interviews1 part · 111 min
Google/Meta/Amazon/Netflix/Uber/Databricks/Snowflake/Microsoft/OpenAI(ML infra)/NVIDIA style mixed rounds: behavioral + system design + scenario + follow-up chains
- Company-Style Interviews111 min
Reviews
to review this kit once you have finished it.