GregLab | Exam Prep

Design and implement an MLOps infrastructure

Azure Machine Learning Workspace Resources

Core

Stand up a workspace with the right dependent resources, connect it to storage through datastores, choose compute that matches the workload, and grant people and workloads only the access they need.

Aligned to the live AI-300 guide, which publishes no skills-measured date; guide and product behavior verified October 10, 2026.

Why this matters

Every training job, model, and endpoint in Azure Machine Learning runs inside a workspace. A workspace built with the wrong storage access pattern, oversized always-on compute, or broad role assignments becomes the source of cost overruns, credential sprawl, and access incidents that later automation cannot fix.

Must Know

  • A workspace is the top-level resource for jobs, assets, models, and endpoints. It depends on an Azure Storage account, an Azure Key Vault, and Application Insights, which are created automatically if you do not bring your own; Azure Container Registry is provisioned when you first build a custom image.
  • A datastore stores connection information to an existing Azure storage service, such as Blob Storage, Azure Data Lake Storage Gen2, or Azure Files. It does not copy data into the workspace.
  • Credential-based datastores authenticate with an account key, SAS token, or service principal, and users with Reader access to the workspace can read those credentials. Identity-based datastores store no secret and use the Microsoft Entra identity of the user or a managed identity; Azure Files datastores support credential-based access only.
  • A compute instance is a single managed VM for interactive development. Schedule it or set idle shutdown so it does not run when nobody is working.
  • A compute cluster scales between a minimum and maximum node count. Setting the minimum to zero lets it release all nodes when no jobs are queued. Low-priority VMs were retired on March 31, 2026; clusters now use Spot VMs for interruptible capacity.
  • Serverless compute runs a job on Microsoft-managed capacity without you creating or managing a cluster. When a job specifies no compute, it runs on serverless compute.
  • Attached Kubernetes compute lets jobs and online inference run on an existing Azure Kubernetes Service or Azure Arc-enabled cluster that you operate.
  • Workspace access uses Azure RBAC. AzureML Data Scientist allows nearly all workspace actions except creating or deleting compute and changing workspace settings; AzureML Compute Operator manages compute.

Compare and Distinguish

  • Datastore versus data asset: the datastore is the connection to storage; a data asset is a named, versioned reference to specific files, folders, or tables reachable through it.
  • Compute instance versus compute cluster: one interactive VM per user versus an autoscaling pool for submitted jobs.
  • Serverless compute versus compute cluster: no cluster object to size and maintain versus a cluster you configure, share, and secure yourself.
  • Workspace managed identity versus user identity: the workspace identity reaches dependent resources; user or compute identities are what an identity-based datastore authorizes for data access.

Scenario examples

  • Scenario: A team wants nightly training to cost nothing when no jobs are queued. Think: a compute cluster with a minimum of zero nodes, or serverless compute.
  • Scenario: Security forbids storing storage account keys anywhere in the workspace. Think: an identity-based datastore with data roles granted to users or compute identities.
  • Scenario: Data scientists must submit jobs but must not create new GPU compute. Think: AzureML Data Scientist for the scientists and AzureML Compute Operator only for the platform team.

Exam traps

  • Registering a datastore does not move or duplicate data; deleting the datastore leaves the underlying storage intact.
  • A compute instance does not autoscale and is not a shared training pool; use a cluster or serverless compute for submitted jobs.
  • Owner or Contributor on the workspace is broader than a data scientist needs; built-in AzureML roles exist for narrower duties.
  • An identity-based datastore still fails if the calling identity lacks a data-plane role such as Storage Blob Data Reader on the storage account.
  • Stopping a compute instance stops compute-hour billing, but disk and networking charges continue.

Key takeaways

  • Connect to data with datastores and prefer identity-based access over stored keys.
  • Match compute to the workload: interactive VM, autoscaling cluster, serverless, or attached Kubernetes.
  • Grant workspace roles by duty, separating job submission from compute administration.
How it works
  • Jobs read data through datastore URIs, so the storage service and its access model decide who can actually read the bytes.
  • Clusters request nodes when jobs queue and release them after the idle period, down to the configured minimum.
Objects and administrative surfaces
  • Azure Machine Learning studio for workspace assets, compute, datastores, and job history.
  • Azure CLI ml extension (v2) commands such as az ml workspace, az ml datastore, and az ml compute with YAML definitions.
  • Azure portal access control (IAM) on the workspace, resource group, or subscription for role assignments.
When to use it
  • Use a compute instance for notebooks and debugging, and a cluster or serverless compute for repeatable training jobs.
  • Use attached Kubernetes when an organization already operates a cluster and wants training or inference to run there.
Security and governance implications
  • Keep secrets in the workspace Key Vault and avoid putting storage keys in code or job definitions.
  • Review role assignments regularly and prefer group-based assignments at workspace or resource-group scope.
Troubleshooting signals
  • A job that fails to read data from an identity-based datastore usually needs a Storage Blob Data role for the user or compute identity.
  • Jobs that queue indefinitely often point to a cluster at its maximum node count or exhausted VM-family quota.
More detail
  • Create a workspace and recognize its dependent resources.
  • Register Blob, Data Lake Storage Gen2, or Azure Files datastores with credential-based or identity-based access.
  • Select compute instances, clusters, serverless compute, or attached Kubernetes for a workload.
  • Assign built-in or custom roles at the narrowest workable scope.

Ready for the quiz?

  • Which Azure resources does a workspace depend on?
  • What changes when a datastore is identity-based rather than credential-based?
  • Why would you set a compute cluster minimum node count to zero?
  • Which built-in role lets someone run experiments without creating compute?

Related objectives

  • D1.1.S1 — Create and manage a workspace
  • D1.1.S2 — Create and manage datastores
  • D1.1.S3 — Create and manage compute targets
  • D1.1.S4 — Configure identity and access management for workspaces

Learn more

Free Microsoft Certified: Machine Learning Operations Engineer Associate prep

Build focused AI-300 quizzes from skill areas, topics, and product references.

Practice with exam-style multiple-choice and multiple-response questions, score breakdowns, explanations, and a compact reference for this lane's official exam domains.

Read Topics Build a quiz

Exam Weights

Exam snapshot

AI-300 at a glance

Level
Intermediate / Associate
Duration
120 minutes
Questions
No fixed live question count published
Formats
No guaranteed question-type mix; the proctored exam may include interactive components
Scoring
Scaled score; 700 minimum passing score

Quiz builder

Choose your practice set

Mode

Exam fidelity: Microsoft does not publish a fixed live question count or guarantee a question-type mix for AI-300. This lane contains multiple-choice and multiple-response exam-style practice. Practice percentages do not reproduce Microsoft's scaled scoring, and difficulty labels describe this site's Intermediate Associate-level MLOps and GenAIOps complexity rather than a Microsoft-published question rating.

Reference

AI-300 topics and reference map

Study links

AI-300 resources