GregLab | Exam Prep

Design and implement a GenAIOps infrastructure

Prompt Design, Variants, and Version Control

Important

Design prompts with clear instructions and grounding, compare prompt variants on the same evaluation data, and keep prompts and agent definitions versioned in Git with evaluation gates on pull requests.

Aligned to the live AI-300 guide, which publishes no skills-measured date; guide and product behavior verified October 10, 2026.

Why this matters

A prompt is production logic. Changing one sentence can shift quality, safety, and cost, so prompts need the same discipline as code: deliberate design, measured comparison, review, and the ability to roll back.

Must Know

  • Effective prompts state the task, provide context and grounding data, specify the output format, and give the model an out, such as saying when the answer is not in the provided data.
  • Few-shot examples and chain-of-thought prompting help non-reasoning models; Microsoft advises that these techniques are not recommended for reasoning models such as gpt-5 and the o-series.
  • Models can show recency bias, so the order and placement of instructions and examples matter.
  • Compare prompt variants by running each against the same test dataset with the same evaluators, and change one thing at a time so the difference is attributable.
  • Prompt flow variants exist only in the Foundry (classic) portal for hub-based projects. Prompt flow is no longer recommended for new development and retires on April 20, 2027, with Microsoft Agent Framework as the migration path.
  • Saved agent versions in Foundry are immutable and referenced as name:version, so a deployment can pin a version and compare definitions.
  • Keep prompt templates and agent definitions as files in a Git repository, review changes through pull requests, and run evaluations as a quality gate before merging.

Compare and Distinguish

  • System instructions versus few-shot examples: rules and role versus demonstrations of the desired output.
  • Prompt variant comparison versus production A/B exposure: offline evaluation on shared data versus live traffic.
  • Git history versus immutable agent versions: source review and rollback of files versus a pinned runtime definition.
  • Prompt flow versus Microsoft Agent Framework: retiring visual flow tool versus code-first framework.

Scenario examples

  • Scenario: Two candidate system prompts must be compared for groundedness. Think: evaluate both on the same dataset with the same evaluators and change only the prompt.
  • Scenario: A prompt change caused off-topic answers in production. Think: revert the commit or redeploy the previous pinned agent version.
  • Scenario: A team plans a new prompt flow project. Think: prompt flow is retiring, so build with Microsoft Agent Framework instead.

Exam traps

  • Editing a prompt directly in production without version control removes review and rollback.
  • Comparing variants on different datasets makes score differences meaningless.
  • Adding chain-of-thought instructions to a reasoning model is not the recommended optimization.
  • Prompt flow variants are not available for Foundry projects in the current portal.

Key takeaways

  • Design prompts with explicit instructions, grounding, format, and an out.
  • Compare variants on identical data and evaluators, one change at a time.
  • Version prompts in Git and gate merges with evaluations.
How it works
  • An evaluation run scores each variant’s outputs with the same evaluators so aggregate metrics can be compared.
  • A pull request triggers a pipeline that runs evaluations and blocks the merge if scores fall below thresholds.
Objects and administrative surfaces
  • Prompt template files and agent definitions stored in a Git repository.
  • Foundry agent versions and evaluation runs that compare versions.
  • GitHub Actions or Azure Pipelines jobs that run evaluations on pull requests.
When to use it
  • Use few-shot examples when a non-reasoning model must follow a precise output format.
  • Use pull request evaluation gates for prompts that affect customer-facing behavior.
Security and governance implications
  • Require reviews for prompt changes in protected branches.
  • Include safety evaluations in prompt change gates, not only quality metrics.
Troubleshooting signals
  • Inconsistent output formats often improve with explicit format instructions or examples placed near the end of the prompt.
  • A variant that wins on one metric but loses on safety should not be promoted without review.
More detail
  • Design and develop prompts with instructions, context, and output formats.
  • Create prompt variants and compare their performance.
  • Implement Git-based version control for prompts and agent definitions.

Ready for the quiz?

  • Why should a prompt give the model an out?
  • What makes a prompt variant comparison fair?
  • Where do prompt flow variants still exist, and when does prompt flow retire?
  • How do you roll back a bad prompt change?

Related objectives

  • D3.3.S1 — Design and develop prompts
  • D3.3.S2 — Create prompt variants and compare performance across different prompts
  • D3.3.S3 — Implement version control for prompts by using Git repositories

Learn more

Free Microsoft Certified: Machine Learning Operations Engineer Associate prep

Build focused AI-300 quizzes from skill areas, topics, and product references.

Practice with exam-style multiple-choice and multiple-response questions, score breakdowns, explanations, and a compact reference for this lane's official exam domains.

Read Topics Build a quiz

Exam Weights

Exam snapshot

AI-300 at a glance

Level
Intermediate / Associate
Duration
120 minutes
Questions
No fixed live question count published
Formats
No guaranteed question-type mix; the proctored exam may include interactive components
Scoring
Scaled score; 700 minimum passing score

Quiz builder

Choose your practice set

Mode

Exam fidelity: Microsoft does not publish a fixed live question count or guarantee a question-type mix for AI-300. This lane contains multiple-choice and multiple-response exam-style practice. Practice percentages do not reproduce Microsoft's scaled scoring, and difficulty labels describe this site's Intermediate Associate-level MLOps and GenAIOps complexity rather than a Microsoft-published question rating.

Reference

AI-300 topics and reference map

Study links

AI-300 resources