GregLab | Exam Prep

Design and implement a GenAIOps infrastructure

Foundation Model Deployment and Capacity

Core

Deploy foundation models as serverless API deployments or on managed compute, choose deployment types for residency, cost, and throughput, manage model versions and retirements, and size provisioned throughput units for high-volume workloads.

Aligned to the live AI-300 guide, which publishes no skills-measured date; guide and product behavior verified October 10, 2026.

Why this matters

Model deployments drive latency, cost, data residency, and outage risk. A deployment type that ignores residency, a version policy that lets a model silently change or stop, or provisioned capacity that is never monitored can break an otherwise healthy generative AI application.

Must Know

  • Serverless API deployments cover Azure OpenAI models and selected partner models such as Anthropic, Mistral, Cohere, and Meta, billed by usage. Open-source and custom models, such as the Hugging Face collection, deploy to managed compute, billed hourly per accelerator and currently in preview.
  • Deployment types apply to serverless API deployments: Global Standard, Data Zone Standard, and Standard; Global, Data Zone, and Regional Provisioned; Global and Data Zone Batch; and Developer for fine-tuned model evaluation.
  • Global Standard is the recommended starting point for most workloads. Data Zone types keep processing within a Microsoft-defined data zone such as the US or EU, and Batch offers a 24-hour target turnaround at lower cost with separate quota.
  • Version upgrade policies are OnceNewDefaultVersionAvailable, OnceCurrentVersionExpired, and NoAutoUpgrade. With NoAutoUpgrade, the deployment stops working at retirement, and provisioned deployments are never auto-upgraded.
  • Retired models return HTTP 410 Gone. Microsoft gives at least 60 days of notice for generally available models and does not grant exceptions that extend a retirement date.
  • Provisioned throughput units (PTUs) reserve model processing capacity for predictable latency. Each model has a minimum deployment size and increment, and quota does not guarantee capacity is available.
  • When provisioned utilization reaches 100 percent, the service returns HTTP 429 with retry-after headers immediately. Spillover can route overflow to a standard deployment in the same resource.
  • Model router selects a model per prompt in Balanced, Quality, or Cost mode, and the model catalog and leaderboards compare models on quality, safety, cost, and throughput.

Compare and Distinguish

  • Serverless API versus managed compute: per-usage billing on Microsoft-hosted capacity versus hourly accelerator billing for open or custom models.
  • Standard versus Provisioned versus Batch: pay as you go versus reserved throughput versus asynchronous discounted processing.
  • Global versus Data Zone versus Standard or Regional Provisioned: processing in any Azure region versus within a Microsoft-defined data zone versus within the deployment's Azure geography.
  • OnceNewDefaultVersionAvailable versus NoAutoUpgrade: early automatic upgrades versus a pinned version that stops at retirement.

Scenario examples

  • Scenario: An EU insurer must keep inference processing inside the EU while using pay-as-you-go pricing. Think: a Data Zone Standard deployment in an EU region.
  • Scenario: A nightly job summarizes millions of documents with no latency requirement. Think: a Global Batch deployment.
  • Scenario: A customer-facing assistant needs consistent latency at very high volume. Think: provisioned throughput sized with the capacity calculator, plus spillover for bursts.

Exam traps

  • Buying a PTU reservation does not guarantee capacity; create the deployment first, then purchase the reservation.
  • A 429 from a provisioned deployment signals full utilization, not a service fault.
  • Managed compute deployments do not use the serverless deployment types.
  • Pinning a version with NoAutoUpgrade avoids surprise changes but causes an outage at retirement if nobody migrates.

Key takeaways

  • Choose the deployment option by model family, then the type by residency, cost, and throughput.
  • Set version upgrade policies deliberately and track retirement dates.
  • Size and monitor PTUs, and plan for 429 handling and spillover.
How it works
  • Requests to a provisioned deployment consume capacity, and utilization is measured continuously against the deployed PTUs.
  • Automatic upgrades replace the model version behind a deployment name according to its upgrade policy.
Objects and administrative surfaces
  • Foundry portal model catalog, leaderboards, and deployments pages.
  • Azure Monitor metric Provisioned-managed Utilization V2 and the capacity calculator for PTU sizing.
  • Model deployment definitions in Bicep or the REST API, including the version upgrade option.
When to use it
  • Use provisioned deployments for steady high volume, standard deployments for variable traffic, and batch for offline jobs.
  • Use managed compute when you need an open-weight model or a custom model that is not offered as a serverless API.
Security and governance implications
  • Use Azure Policy and deployment type choices to enforce residency requirements.
  • Review model retirement notices regularly and test replacement versions before the date.
Troubleshooting signals
  • A deployment returning 410 Gone is calling a retired model version and must move to a supported version.
  • Frequent 429 responses with high utilization indicate undersized PTUs or missing spillover configuration.
More detail
  • Deploy models through serverless API deployments and managed compute.
  • Select models with the catalog, leaderboards, and model router.
  • Implement version upgrade policies and retirement plans.
  • Size, purchase, and monitor provisioned throughput units.

Ready for the quiz?

  • Which models deploy to managed compute instead of serverless API deployments?
  • Which deployment type keeps processing in a data zone with pay-as-you-go billing?
  • What happens to a NoAutoUpgrade deployment at retirement?
  • How should a client respond to a 429 from a provisioned deployment?

Related objectives

  • D3.2.S1 — Deploy foundation models by using serverless API endpoints and managed compute options
  • D3.2.S2 — Select appropriate models for specific use cases
  • D3.2.S3 — Implement model versioning and production deployment strategies
  • D3.2.S4 — Configure provisioned throughput units for high-volume workloads

Learn more

Free Microsoft Certified: Machine Learning Operations Engineer Associate prep

Build focused AI-300 quizzes from skill areas, topics, and product references.

Practice with exam-style multiple-choice and multiple-response questions, score breakdowns, explanations, and a compact reference for this lane's official exam domains.

Read Topics Build a quiz

Exam Weights

Exam snapshot

AI-300 at a glance

Level
Intermediate / Associate
Duration
120 minutes
Questions
No fixed live question count published
Formats
No guaranteed question-type mix; the proctored exam may include interactive components
Scoring
Scaled score; 700 minimum passing score

Quiz builder

Choose your practice set

Mode

Exam fidelity: Microsoft does not publish a fixed live question count or guarantee a question-type mix for AI-300. This lane contains multiple-choice and multiple-response exam-style practice. Practice percentages do not reproduce Microsoft's scaled scoring, and difficulty labels describe this site's Intermediate Associate-level MLOps and GenAIOps complexity rather than a Microsoft-published question rating.

Reference

AI-300 topics and reference map

Study links

AI-300 resources