GregLab | Exam Prep

Implement generative AI quality assurance and observability

Generative AI Evaluation and Validation

Core

Build test datasets and field mappings, measure groundedness, relevance, coherence, and fluency, run risk and safety evaluators and red teaming, and automate evaluations with built-in and custom evaluators in CI.

Aligned to the live AI-300 guide, which publishes no skills-measured date; guide and product behavior verified October 10, 2026.

Why this matters

Generative AI output cannot be validated by exact-match tests alone. Evaluators give measurable quality and safety signals, and automating them in pipelines stops regressions before users see them.

Must Know

  • Evaluation datasets are typically JSONL files with one JSON object per row. Each row supplies fields such as query, response, context, and ground_truth, and mappings tell each evaluator which field to read; current Foundry evaluations use {{item.query}} syntax, while the classic SDK used ${data.query}.
  • Groundedness checks whether the response is supported by the provided context and needs response and context. Relevance checks whether the response addresses the query and needs query and response.
  • Coherence measures logical, well-organized responses, and fluency measures language quality. These AI-assisted quality evaluators score 1 to 5 with a default pass threshold of 3.
  • Similarity needs ground truth and a judge deployment, while F1, BLEU, GLEU, ROUGE, and METEOR are 0 to 1 overlap scores computed without a model deployment.
  • Risk and safety evaluators cover violence, sexual, self-harm, and hate and unfairness content on a 0 to 7 severity scale, plus indirect attack, protected material, and code vulnerability. They do not need your own model deployment.
  • Agent evaluators include intent resolution, tool call accuracy, and task adherence; tool call accuracy needs the tool definitions.
  • The AI Red Teaming Agent uses PyRIT attack strategies and reports attack success rate.
  • Custom evaluators can be code-based or prompt-based. Automated evaluation can run in GitHub Actions or Azure DevOps, and Microsoft advises against running it on every commit to control cost.

Compare and Distinguish

  • Groundedness versus relevance: supported by context versus responsive to the question.
  • Coherence versus fluency: logical flow of ideas versus grammatical, natural language.
  • AI-assisted quality evaluators versus overlap metrics: judge-model scoring versus string comparison with ground truth.
  • Content safety evaluators versus red teaming: scoring responses for harm versus actively probing with attacks.

Scenario examples

  • Scenario: A support bot sometimes states policy details that are not in the retrieved documents. Think: groundedness with response and context.
  • Scenario: Answers are accurate but often ignore what the user asked. Think: relevance.
  • Scenario: Documents retrieved from the web might contain hidden instructions. Think: the indirect attack evaluator and red teaming.

Exam traps

  • Fluency does not measure factual accuracy or grounding.
  • Similarity and overlap metrics need ground truth; groundedness needs context instead.
  • A safety severity score passes when it is at or below the threshold, unlike quality scores where higher is better.
  • Running evaluation on every commit increases cost without improving gating.

Key takeaways

  • Map the right dataset fields to each evaluator.
  • Pick evaluators by the failure you need to detect, quality or safety.
  • Automate evaluations at pull request or release points with thresholds.
How it works
  • AI-assisted evaluators prompt a judge model with the mapped fields and return a score and reason.
  • CI evaluation compares versions with confidence intervals and statistical significance tests.
Objects and administrative surfaces
  • Foundry portal evaluations, evaluator catalog, and results views.
  • Evaluation SDK runs over JSONL datasets with field mappings.
  • GitHub Actions and Azure DevOps tasks that run agent evaluations in pipelines.
When to use it
  • Use overlap metrics when exact reference answers exist, and AI-assisted metrics for open-ended responses.
  • Use red teaming before release and after major prompt or model changes.
Security and governance implications
  • Treat safety evaluation results as release criteria alongside quality.
  • Keep evaluation datasets free of sensitive data or protect them like production data.
Troubleshooting signals
  • An evaluator that fails with missing inputs usually has an incorrect field mapping.
  • Unexpectedly low groundedness can come from mapping the wrong context field rather than from the model.
More detail
  • Create test datasets and map their fields to evaluators.
  • Implement groundedness, relevance, coherence, and fluency metrics.
  • Configure risk and safety evaluations and red teaming.
  • Automate evaluation workflows with built-in and custom evaluators.

Ready for the quiz?

  • Which fields does groundedness need?
  • How does coherence differ from fluency?
  • Which evaluators need no judge-model deployment?
  • What does attack success rate report?

Related objectives

  • D4.1.S1 — Create test datasets and data mapping for comprehensive model evaluation
  • D4.1.S2 — Implement AI quality metrics, including groundedness, relevance, coherence, and fluency
  • D4.1.S3 — Configure risk and safety evaluations for harmful content detection
  • D4.1.S4 — Set up automated evaluation workflows by using built-in and custom evaluation metrics

Learn more

Free Microsoft Certified: Machine Learning Operations Engineer Associate prep

Build focused AI-300 quizzes from skill areas, topics, and product references.

Practice with exam-style multiple-choice and multiple-response questions, score breakdowns, explanations, and a compact reference for this lane's official exam domains.

Read Topics Build a quiz

Exam Weights

Exam snapshot

AI-300 at a glance

Level
Intermediate / Associate
Duration
120 minutes
Questions
No fixed live question count published
Formats
No guaranteed question-type mix; the proctored exam may include interactive components
Scoring
Scaled score; 700 minimum passing score

Quiz builder

Choose your practice set

Mode

Exam fidelity: Microsoft does not publish a fixed live question count or guarantee a question-type mix for AI-300. This lane contains multiple-choice and multiple-response exam-style practice. Practice percentages do not reproduce Microsoft's scaled scoring, and difficulty labels describe this site's Intermediate Associate-level MLOps and GenAIOps complexity rather than a Microsoft-published question rating.

Reference

AI-300 topics and reference map

Study links

AI-300 resources