Intermediate

MLOps Engineer (Classical)

625 interview questions for MLOps engineers, answered

14 parts20h 36mRevised Oct 2026

Overview

The classical MLOps engineer interview, answered. Fourteen volumes cover the production mindset, data engineering for MLOps, experiment tracking and the model lifecycle, testing ML systems, CI/CD and continuous training, model serving and inference, monitoring and reliability, model governance, ML security, cloud MLOps, performance and cost engineering, production ML system design, debugging and failure diagnosis, and company-style interview rounds. Every answer is written in the first person with a concrete example, and many carry a diagram.

What you will learn

  • Explain why MLOps exists and what it adds to DevOps
  • Build data pipelines and track experiments and models
  • Test ML systems and automate CI/CD and continuous training
  • Serve models and monitor them for drift and failures
  • Govern and secure models in production
  • Run MLOps on the major clouds and control cost
  • Design production ML systems and debug them under pressure
  • Handle company-style MLOps interview rounds

Included with the kit

  • 14 written parts, yours for good
  • 20h 36m of reading, measured not estimated
  • Written for the intermediate level
  • Every future revision included
  • UPI, cards and netbanking

Prerequisites

  • Software engineering or data engineering experience
  • Basic machine learning concepts

Curriculum

14 sections · 14 parts · 20h 36m
01Volume 1: Foundations & Production Mindset1 part · 82 min

ML lifecycle, CRISP-DM, DevOps vs DataOps vs MLOps vs LLMOps, ML platform evolution, maturity levels, orchestration (Airflow/Kubeflow/Dagster/Prefect), reproducibility, env mgmt (Docker/Conda/Poetry), Git/GitOps, Terraform/IaC, secrets mgmt, repo structure

  • Foundations & Production Mindset82 min
02Volume 2: Data Engineering for MLOps1 part · 106 min

Ingestion (batch/streaming/Kafka/Spark), Delta Lake/Iceberg, DVC/LakeFS, data contracts, lineage, schema evolution, feature engineering, feature stores (Feast/Tecton), online vs offline, point-in-time correctness, data validation (Great Expectations/TFDV/Evidently), feature freshness, drift sources

  • Data Engineering for MLOps106 min
03Volume 3: Experiment Tracking & Model Lifecycle1 part · 82 min

MLflow/W&B/Comet/Neptune, artifacts, lineage, metadata, registry, champion/challenger, approval, promotion, rollback, reproducibility, versioning, auditability

  • Experiment Tracking & Model Lifecycle82 min
04Volume 4: Testing ML Systems1 part · 81 min

Unit/data/feature/pipeline/integration/regression/shadow/smoke/acceptance testing, golden datasets, model validation, CI testing

  • Testing ML Systems81 min
05Volume 5: CI/CD/CT for ML1 part · 58 min

GitHub Actions/GitLab CI/Azure DevOps/Jenkins, continuous training/evaluation/deployment/delivery, model promotion, rollback automation

  • CI/CD/CT for ML58 min
06Volume 6: Model Serving & Inference1 part · 68 min

FastAPI/gRPC/REST, Docker/K8s/Helm, KEDA/HPA, GPU scheduling, BentoML/Triton/Seldon/TorchServe, canary/blue-green/shadow, autoscaling, ONNX/TensorRT/OpenVINO, quantization, batching, rate limiting, caching, auth

  • Model Serving & Inference68 min
07Volume 7: Monitoring & Reliability1 part · 85 min

Prometheus/Grafana/Datadog/OpenTelemetry, logging/tracing/metrics, latency/CPU/GPU/memory, business metrics, concept/prediction/feature drift, training-serving skew, alerting, SHAP/LIME, fairness/bias, RCA

  • Monitoring & Reliability85 min
08Volume 8: Model Governance1 part · 38 min

Compliance, lineage, auditability, model cards, approval, RBAC, version control, access control, PII, data retention, documentation

  • Model Governance38 min
09Volume 9: ML Security1 part · 83 min

Model extraction, data/model poisoning, adversarial examples, secrets, API abuse, dependency scanning, container security, supply chain. Prompt injection is compared briefly here and covered in full in the AI Security kit.

  • ML Security83 min
10Volume 10: Cloud MLOps1 part · 87 min

AWS (SageMaker/S3/Lambda/EKS/ECS/ECR), Azure (Azure ML/AKS/Blob), GCP (Vertex AI/Cloud Run/BigQuery/GKE), multi-cloud, hybrid cloud

  • Cloud MLOps87 min
11Volume 11: Performance & Cost Engineering1 part · 76 min

GPU/CPU optimization, autoscaling, compression/distillation/quantization, batching, caching, spot instances, scheduling, resource allocation, profiling, benchmarking

  • Performance & Cost Engineering76 min
12Volume 12: Production ML System Design1 part · 104 min

Recsys, fraud detection, search, ads, ranking, OCR, medical AI, forecasting, vision, speech, real-time/streaming ML, multi-region, HA, disaster recovery

  • Production ML System Design104 min
13Volume 13: Debugging & Failure Diagnosis1 part · 175 min

Broken pipelines, OOM, CrashLoopBackOff, feature drift, registry failures, deployment failures, timeouts, schema mismatch, version conflicts, training-serving mismatch, silent failures, RCA

  • Debugging & Failure Diagnosis175 min
14Volume 14: Company-Style Interviews1 part · 111 min

Google/Meta/Amazon/Netflix/Uber/Databricks/Snowflake/Microsoft/OpenAI(ML infra)/NVIDIA style mixed rounds: behavioral + system design + scenario + follow-up chains

  • Company-Style Interviews111 min

Reviews

No reviews yet. The first comes from a reader who finishes the kit.

to review this kit once you have finished it.

Next to this one