Implement machine learning model lifecycle and operations
Model Training Orchestration
CoreTrack experiments with MLflow, explore candidates with automated ML and notebooks, run reproducible command jobs, tune hyperparameters with sweep jobs, scale out with distributed training, and chain steps into pipelines.
Aligned to the live AI-300 guide, which publishes no skills-measured date; guide and product behavior verified October 10, 2026.
Why this matters
Training becomes operational only when it is repeatable, comparable, and automated. The decisions here determine whether a team can explain why one model beat another, rerun last month’s training exactly, and scale to large models without rewriting everything.
Must Know
- Azure Machine Learning workspaces are MLflow-compatible, and the SDK v2 has no logging API of its own. On Azure compute the tracking URI is configured automatically, and mlflow.autolog() captures parameters, metrics, and models for supported frameworks.
- To log from a local machine or another environment, set the MLflow tracking URI to the workspace tracking URI before logging.
- Automated ML tries algorithms and featurization for classification, regression, forecasting, computer vision, and NLP tasks, ranked by the primary metric you choose; tabular tasks take mltable input.
- A command job runs a script with a specified code folder, command, environment, compute, and inputs and outputs, which makes notebook experiments reproducible.
- A sweep job tunes hyperparameters over a search space with a sampling algorithm (random, grid, or Bayesian), a primary metric and goal, and trial limits.
- Early termination policies (bandit, median stopping, truncation selection) stop poorly performing trials. Without a policy, every trial runs to completion. Random and grid sampling support early termination; grid sampling accepts only choice values.
- Distributed training sets instance_count for the number of nodes and process_count_per_instance for processes per node, commonly one per GPU, with PyTorch, TensorFlow, MPI, or DeepSpeed.
- Pipeline jobs connect components through inputs and outputs, can reuse results of unchanged steps, and can run on a schedule.
- Compare jobs in studio by selecting runs to chart metrics side by side, or query runs programmatically with MLflow.
Compare and Distinguish
- Automated ML versus sweep job: Azure Machine Learning chooses algorithms and featurization versus you supply the training script and tune its hyperparameters.
- Random versus grid versus Bayesian sampling: broad cheap exploration versus exhaustive discrete combinations versus sequential choices informed by prior trials.
- Bandit versus median stopping versus truncation selection: slack against the best run versus below the running median versus cutting a fixed percentage of the worst runs.
- Command job versus pipeline job: one script execution versus a graph of reusable steps.
Scenario examples
- Scenario: A tuning run wastes budget on trials that are clearly behind after a few epochs. Think: add a bandit or median stopping policy to a random-sampling sweep.
- Scenario: A model is trained on four nodes with eight GPUs each. Think: instance_count of four and process_count_per_instance of eight.
- Scenario: A data scientist logged metrics from a laptop but nothing appears in the workspace. Think: set the MLflow tracking URI to the workspace.
Exam traps
- The sweep primary_metric must exactly match the metric name the training script logs, or trials cannot be compared.
- Grid sampling only works with discrete choice values, not continuous distributions.
- process_count_per_instance is per node; setting it to the total GPU count across nodes oversubscribes each node.
- A notebook run is not reproducible automation; convert it to a script and command job.
Key takeaways
- Log everything with MLflow and compare runs on the primary metric.
- Use automated ML to explore and sweep jobs to tune your own script.
- Scale with distributed settings and automate with pipelines and schedules.
How it works
- Each sweep trial is a child job with its own parameters and metrics, evaluated against the primary metric.
- Pipeline steps run when their inputs are ready, and unchanged steps can reuse previous outputs.
Objects and administrative surfaces
- Studio Jobs pages for metrics, logs, outputs, and run comparison.
- Python SDK v2 command(), sweep(), and pipeline definitions or equivalent CLI YAML.
- MLflow APIs such as mlflow.autolog, mlflow.log_metric, and mlflow.search_runs.
When to use it
- Use DeepSpeed or other memory-optimizing approaches when a large model does not fit on a single GPU.
- Use pipeline schedules for recurring retraining on fresh data.
Security and governance implications
- Run training on compute with managed identities so scripts do not embed credentials.
- Keep training scripts and job YAML in Git so tuning changes are reviewable.
Troubleshooting signals
- A sweep that never stops early may have no termination policy or an evaluation interval longer than the trial length.
- Distributed jobs that hang at startup often have mismatched process counts or networking between nodes.
More detail
- Configure MLflow tracking and autologging.
- Run automated ML with a suitable task, primary metric, and limits.
- Use notebooks for exploration and command jobs for reproducible training.
- Configure sweep jobs, sampling, and early termination.
- Configure distributed training and pipelines, then compare jobs.
Ready for the quiz?
- Which sampling algorithm requires choice-only search spaces?
- What do instance_count and process_count_per_instance control?
- When should a team prefer automated ML over a sweep job?
- How do you make local MLflow logging land in the workspace?
Related objectives
- D2.1.S1 — Configure experiment tracking with MLflow
- D2.1.S2 — Use automated machine learning to explore optimal models
- D2.1.S3 — Use notebooks for experimentation and exploration
- D2.1.S4 — Automate hyperparameter tuning
- D2.1.S5 — Run model training scripts
- D2.1.S6 — Manage distributed training for large and deep learning models
- D2.1.S7 — Implement training pipelines
- D2.1.S8 — Compare model performance across jobs