Generative AI platform
Amazon Bedrock Model Evaluation
Compares foundation model outputs with automatic and human evaluation workflows.
Key points
- Helps compare model responses before selecting a model for a use case.
- Supports automated evaluation with datasets and metrics, and human evaluation when judgment is needed.
- Can evaluate dimensions such as quality, relevance, toxicity, and task-specific performance.
- Supports informed model selection rather than choosing only by provider or model size.
- Fits governance workflows where teams need evidence for model approval.
When to use it
- Choose Model Evaluation when several foundation models must be compared for a task.
- Use it before moving a generative AI application into production.
- Use human evaluation when the desired output quality depends on subjective business judgment.
Exam tips
- Evaluation selects and validates models; Guardrails enforce runtime safety controls.
- Human evaluation is useful when automatic metrics cannot capture usefulness or brand fit.
- Do not confuse model evaluation with SageMaker Clarify, which focuses on bias and explainability for ML models.
- Evaluation results guide model choice but do not automatically create a tuned model.