Design and implement a GenAIOps infrastructure
Prompt Design, Variants, and Version Control
ImportantDesign prompts with clear instructions and grounding, compare prompt variants on the same evaluation data, and keep prompts and agent definitions versioned in Git with evaluation gates on pull requests.
Aligned to the live AI-300 guide, which publishes no skills-measured date; guide and product behavior verified October 10, 2026.
Why this matters
A prompt is production logic. Changing one sentence can shift quality, safety, and cost, so prompts need the same discipline as code: deliberate design, measured comparison, review, and the ability to roll back.
Must Know
- Effective prompts state the task, provide context and grounding data, specify the output format, and give the model an out, such as saying when the answer is not in the provided data.
- Few-shot examples and chain-of-thought prompting help non-reasoning models; Microsoft advises that these techniques are not recommended for reasoning models such as gpt-5 and the o-series.
- Models can show recency bias, so the order and placement of instructions and examples matter.
- Compare prompt variants by running each against the same test dataset with the same evaluators, and change one thing at a time so the difference is attributable.
- Prompt flow variants exist only in the Foundry (classic) portal for hub-based projects. Prompt flow is no longer recommended for new development and retires on April 20, 2027, with Microsoft Agent Framework as the migration path.
- Saved agent versions in Foundry are immutable and referenced as name:version, so a deployment can pin a version and compare definitions.
- Keep prompt templates and agent definitions as files in a Git repository, review changes through pull requests, and run evaluations as a quality gate before merging.
Compare and Distinguish
- System instructions versus few-shot examples: rules and role versus demonstrations of the desired output.
- Prompt variant comparison versus production A/B exposure: offline evaluation on shared data versus live traffic.
- Git history versus immutable agent versions: source review and rollback of files versus a pinned runtime definition.
- Prompt flow versus Microsoft Agent Framework: retiring visual flow tool versus code-first framework.
Scenario examples
- Scenario: Two candidate system prompts must be compared for groundedness. Think: evaluate both on the same dataset with the same evaluators and change only the prompt.
- Scenario: A prompt change caused off-topic answers in production. Think: revert the commit or redeploy the previous pinned agent version.
- Scenario: A team plans a new prompt flow project. Think: prompt flow is retiring, so build with Microsoft Agent Framework instead.
Exam traps
- Editing a prompt directly in production without version control removes review and rollback.
- Comparing variants on different datasets makes score differences meaningless.
- Adding chain-of-thought instructions to a reasoning model is not the recommended optimization.
- Prompt flow variants are not available for Foundry projects in the current portal.
Key takeaways
- Design prompts with explicit instructions, grounding, format, and an out.
- Compare variants on identical data and evaluators, one change at a time.
- Version prompts in Git and gate merges with evaluations.
How it works
- An evaluation run scores each variant’s outputs with the same evaluators so aggregate metrics can be compared.
- A pull request triggers a pipeline that runs evaluations and blocks the merge if scores fall below thresholds.
Objects and administrative surfaces
- Prompt template files and agent definitions stored in a Git repository.
- Foundry agent versions and evaluation runs that compare versions.
- GitHub Actions or Azure Pipelines jobs that run evaluations on pull requests.
When to use it
- Use few-shot examples when a non-reasoning model must follow a precise output format.
- Use pull request evaluation gates for prompts that affect customer-facing behavior.
Security and governance implications
- Require reviews for prompt changes in protected branches.
- Include safety evaluations in prompt change gates, not only quality metrics.
Troubleshooting signals
- Inconsistent output formats often improve with explicit format instructions or examples placed near the end of the prompt.
- A variant that wins on one metric but loses on safety should not be promoted without review.
More detail
- Design and develop prompts with instructions, context, and output formats.
- Create prompt variants and compare their performance.
- Implement Git-based version control for prompts and agent definitions.
Ready for the quiz?
- Why should a prompt give the model an out?
- What makes a prompt variant comparison fair?
- Where do prompt flow variants still exist, and when does prompt flow retire?
- How do you roll back a bad prompt change?
Related objectives
- D3.3.S1 — Design and develop prompts
- D3.3.S2 — Create prompt variants and compare performance across different prompts
- D3.3.S3 — Implement version control for prompts by using Git repositories