Design and implement an MLOps infrastructure
Azure Machine Learning Assets and Registries
CoreVersion the data, environments, and components that jobs depend on, and share proven assets across workspaces through registries so promotion does not mean rebuilding.
Aligned to the live AI-300 guide, which publishes no skills-measured date; guide and product behavior verified October 10, 2026.
Why this matters
Reproducibility depends on knowing exactly which data version, software environment, and pipeline step produced a model. Assets give each of those a name and version, and registries let a platform team promote the same versions from development to production workspaces.
Must Know
- Data assets reference data in storage and come in three types: uri_file for a single file, uri_folder for a folder, and mltable for tabular data with a schema definition.
- Data asset versions are immutable references. Pinning a version in a job makes the input reproducible even when new data versions are registered later.
- An environment captures the software for a job: a Docker image, optionally with a conda specification layered on top, or a Docker build context.
- Curated environments are Microsoft-maintained and ready to use. Create a custom environment when you need packages or versions the curated ones lack.
- A component is a self-contained, versioned step with a defined interface of inputs, outputs, parameters, command, code, and environment. Pipelines are built from components.
- A registry is an organization-wide store that sits outside any single workspace, so workspaces in different regions or subscriptions can use the same models, components, and environments.
- Assets created in a registry can be referenced from any workspace that has access to the registry, which supports dev-to-test-to-prod promotion without copying files by hand.
Compare and Distinguish
- uri_folder versus mltable: a folder path read as files versus a table definition that tells consumers how to load and type the data.
- Curated versus custom environment: Microsoft-maintained images versus your own image or conda file with versions you control.
- Component versus environment: the reusable step and its interface versus the software stack the step runs on.
- Workspace asset versus registry asset: visible inside one workspace versus shareable across many workspaces and regions.
Scenario examples
- Scenario: An automated ML regression job needs typed tabular input. Think: register the training data as an mltable data asset.
- Scenario: A training script needs a package that no curated environment includes. Think: a custom environment from a base image plus a conda file.
- Scenario: Production workspaces in another subscription must deploy the exact model validated in development. Think: register the model in a registry and deploy from the registry reference.
Exam traps
- Registering a data asset does not copy the data; deleting the underlying files breaks every version that points to them.
- The files, data, name, and version number of an existing asset version cannot change, so you register a new version, although description and tags can still be updated.
- Sharing a workspace with another team is not the same as a registry; a registry is designed for cross-workspace reuse and promotion.
- A component without a pinned environment version can produce different results when the environment changes.
Key takeaways
- Pick the data asset type by how the consumer reads the data.
- Pin data, environment, and component versions for reproducible jobs.
- Use registries to promote assets across workspaces, regions, and subscriptions.
How it works
- A job resolves each asset reference to a specific version at submission time and records it in lineage.
- Environment definitions are built into images that are cached and reused when the definition is unchanged.
Objects and administrative surfaces
- Studio Data, Environments, and Components pages for listing versions and lineage.
- YAML definitions for data assets, environments, and components created with az ml data create, az ml environment create, and az ml component create.
- Registries created with az ml registry create and referenced with azureml://registries/ paths.
When to use it
- Use data assets whenever a dataset will be reused or must be traceable to a model.
- Use components when a step will appear in more than one pipeline or be owned by a different team.
Security and governance implications
- Control who can create or read registry assets with Azure RBAC on the registry.
- Keep environment definitions in source control so package changes are reviewed.
Troubleshooting signals
- An environment image build failure usually traces to an unresolvable package version or an unreachable package feed.
- A pipeline that unexpectedly reruns a step often has an input, code, or environment change that invalidates reuse.
More detail
- Create and version uri_file, uri_folder, and mltable data assets.
- Use curated environments and build custom ones from images, conda files, or build contexts.
- Define components with inputs, outputs, code, command, and environment.
- Share models, components, and environments through registries.
Ready for the quiz?
- When is mltable a better choice than uri_folder?
- What three ways can you define a custom environment?
- What does a component interface define?
- Why would a team use a registry instead of re-registering a model in each workspace?
Related objectives
- D1.2.S1 — Create and manage data assets
- D1.2.S2 — Create and manage environments
- D1.2.S3 — Create and manage components
- D1.2.S4 — Share assets across workspaces by using registries