GregLab | Exam Prep

Implement machine learning model lifecycle and operations

Production Model Deployment

Core

Serve models through managed online endpoints, Kubernetes online endpoints, or batch endpoints; debug deployments locally and from logs; and roll new versions out gradually with traffic splits, mirrored traffic, and fast rollback.

Aligned to the live AI-300 guide, which publishes no skills-measured date; guide and product behavior verified October 10, 2026.

Why this matters

Deployment is where a model starts affecting users and costs money every hour. Choosing the right endpoint type, testing safely before shifting traffic, and being able to roll back in seconds are what separate a routine release from an outage.

Must Know

  • An endpoint is a stable URL and authentication boundary; deployments under it host model versions with their own compute, so one endpoint can run several deployments.
  • Managed online endpoints are the recommended online option: Azure manages the serving infrastructure, scaling, and OS updates. Kubernetes online endpoints run on clusters you operate.
  • Batch endpoints process large volumes asynchronously on compute clusters. A model deployment runs a model with optional scoring script, and a pipeline component deployment runs a reusable pipeline.
  • Live traffic allocation across deployments must total 100 percent, or every deployment can be set to 0 percent, in which case the endpoint returns HTTP 404. A request can target a specific deployment with the azureml-model-deployment header, bypassing the split.
  • Mirrored traffic copies a percentage of live requests to a shadow deployment. Its responses are not returned to clients, but its metrics and logs are recorded.
  • Mirroring supports one shadow deployment per endpoint, at most 50 percent of traffic, and is not supported on Kubernetes online endpoints.
  • Local endpoints run a deployment in Docker on your machine with the --local flag, which is faster for debugging scoring scripts but supports only one deployment and no traffic rules or authentication.
  • az ml online-deployment get-logs returns inference server logs; add --container storage-initializer to see model download problems.

Compare and Distinguish

  • Online versus batch endpoint: synchronous low-latency requests versus asynchronous jobs over large datasets.
  • Managed versus Kubernetes online endpoint: Azure-managed infrastructure versus clusters your team provisions and patches.
  • Traffic split versus mirrored traffic: users receive responses from the new deployment versus the new deployment is tested invisibly.
  • Rollback versus redeploy: shifting traffic back to the previous deployment versus rebuilding an old version from scratch.

Scenario examples

  • Scenario: Ten million records must be scored overnight with no latency requirement. Think: a batch endpoint on a compute cluster.
  • Scenario: A new model must be validated against real requests without any customer seeing its output. Think: mirror a small share of traffic to the new deployment.
  • Scenario: The green deployment raises error rates after receiving 20 percent of traffic. Think: set its traffic back to zero so blue serves everything, then investigate.

Exam traps

  • Mirrored traffic is not a canary release; users never see the shadow deployment output.
  • A deployment can receive live traffic or mirrored traffic, not both at once.
  • Requests that name a deployment in the header are not mirrored.
  • Deleting the old deployment immediately after switching traffic removes the fastest rollback path.

Key takeaways

  • Pick online or batch by latency and volume, then managed or Kubernetes by who runs the infrastructure.
  • Test new deployments with header routing or mirroring before shifting live traffic.
  • Roll forward gradually and keep the previous deployment until the new one is proven.
How it works
  • The endpoint routes each request according to traffic weights unless a deployment header overrides it.
  • Each deployment runs one or more instances; Microsoft recommends at least three for high availability and reserves extra capacity for upgrades.
Objects and administrative surfaces
  • az ml online-endpoint and az ml online-deployment commands with endpoint and deployment YAML.
  • az ml batch-endpoint and az ml batch-deployment commands for batch scoring.
  • Studio Endpoints pages for traffic, logs, metrics, and test calls.
When to use it
  • Use batch endpoints for periodic large-scale scoring and online endpoints for interactive applications.
  • Use Kubernetes online endpoints when inference must run on existing or on-premises clusters.
Security and governance implications
  • Use Microsoft Entra token authentication for endpoints where callers have identities, and keys only where required.
  • Batch endpoint jobs run under the identity of the invoker. Inputs from credential-based datastores use the datastore's account key or SAS token, while credential-less inputs are read with the job identity and mounted with the compute cluster managed identity, which needs at least Storage Blob Data Reader.
Troubleshooting signals
  • A deployment stuck in a crash loop usually has an exception in the scoring script init function or a missing package.
  • HTTP 404 from an endpoint can mean no deployment has positive traffic weight; HTTP 429 indicates too many pending requests or rate limiting.
More detail
  • Deploy models to managed online, Kubernetes online, and batch endpoints.
  • Test deployments locally and through header routing, and troubleshoot with logs.
  • Implement blue-green rollout, mirrored traffic, and rollback.

Ready for the quiz?

  • What happens to responses from a mirrored deployment?
  • What are the limits on mirrored traffic?
  • How can you test a new deployment that has zero percent of traffic?
  • Which log container shows model download failures?

Related objectives

  • D2.3.S1 — Deploy models as real-time or batch endpoints with managed inference options
  • D2.3.S2 — Test and troubleshoot model endpoints
  • D2.3.S3 — Implement progressive rollout and safe rollback strategies

Learn more

Free Microsoft Certified: Machine Learning Operations Engineer Associate prep

Build focused AI-300 quizzes from skill areas, topics, and product references.

Practice with exam-style multiple-choice and multiple-response questions, score breakdowns, explanations, and a compact reference for this lane's official exam domains.

Read Topics Build a quiz

Exam Weights

Exam snapshot

AI-300 at a glance

Level
Intermediate / Associate
Duration
120 minutes
Questions
No fixed live question count published
Formats
No guaranteed question-type mix; the proctored exam may include interactive components
Scoring
Scaled score; 700 minimum passing score

Quiz builder

Choose your practice set

Mode

Exam fidelity: Microsoft does not publish a fixed live question count or guarantee a question-type mix for AI-300. This lane contains multiple-choice and multiple-response exam-style practice. Practice percentages do not reproduce Microsoft's scaled scoring, and difficulty labels describe this site's Intermediate Associate-level MLOps and GenAIOps complexity rather than a Microsoft-published question rating.

Reference

AI-300 topics and reference map

Study links

AI-300 resources