Implement generative AI quality assurance and observability
Generative AI Observability and Cost Monitoring
CoreMonitor generative AI applications and agents in Foundry with Application Insights, track latency, throughput, and token-based cost metrics in Azure Monitor, and configure tracing and diagnostic logs for production troubleshooting.
Aligned to the live AI-300 guide, which publishes no skills-measured date; guide and product behavior verified October 10, 2026.
Why this matters
Production generative AI fails in ways that only telemetry reveals: slow first tokens, runaway token use, failing tool calls, and drifting quality. Observability lets operators see each step of a run and connect cost to the deployments and projects that create it.
Must Know
- Connecting an Application Insights resource to a Foundry project enables server-side tracing for agents without code changes, and the agent monitoring dashboard reads its data from that resource.
- Foundry stores traces in Application Insights using OpenTelemetry semantic conventions, so each run, model call, and tool call appears as a span. Querying traces needs Log Analytics Reader.
- Continuous evaluation samples production runs and scores them with evaluators, with a default limit of 100 runs per hour.
- Azure Monitor collects Foundry Models platform metrics automatically. Time to Response is the recommended latency measure where it is available, and Tokens Per Second and Time Between Tokens describe generation speed.
- Token metrics such as Processed Prompt Tokens and Generated Completion Tokens show usage per deployment, and Provisioned-managed Utilization V2 shows PTU saturation.
- Azure Cost Management shows per-deployment cost with a delay of several hours, so token metrics are the faster signal for spikes.
- Diagnostic settings send logs such as RequestResponse, Trace, and AzureOpenAIRequestUsage to a Log Analytics workspace, and they must be configured on each resource.
Compare and Distinguish
- Tracing versus metrics: per-request step detail versus aggregated time series.
- Time to Response versus Tokens Per Second: wait before the first output versus generation speed.
- Continuous evaluation versus offline evaluation: sampled production runs versus fixed test datasets.
- Token metrics versus Cost Management: near-real-time usage versus billed cost after a delay.
Scenario examples
- Scenario: Users complain that the assistant takes too long before it starts answering. Think: Time to Response metrics and traces of the slow requests.
- Scenario: Monthly spend doubled after a release. Think: Processed Prompt Tokens by deployment, then traces to find the prompt growth.
- Scenario: An agent sometimes fails in the middle of a multistep task. Think: OpenTelemetry traces showing which tool call span failed.
Exam traps
- The legacy Latency metric is described as misleading; use Time to Response for responsiveness.
- Platform metrics do not show the content of prompts and responses; tracing and logs are needed for step detail.
- Diagnostic settings applied to one resource do not cover other resources.
- Cost Management data is delayed, so it is not the right real-time alert source.
Key takeaways
- Connect Application Insights to get traces and monitoring dashboards.
- Watch Time to Response, throughput, tokens, and utilization in Azure Monitor.
- Route diagnostic logs to Log Analytics for deep troubleshooting.
How it works
- Instrumentation emits spans for each operation, linked by trace IDs into a single run view.
- Metrics are aggregated per deployment and can be split by dimensions such as model or project.
Objects and administrative surfaces
- Foundry portal monitoring dashboard and trace views backed by Application Insights.
- Azure Monitor metrics explorer and alert rules for Foundry Models.
- Diagnostic settings that route resource logs to Log Analytics.
When to use it
- Use metric alerts for latency and utilization thresholds and traces for root cause.
- Use continuous evaluation when production quality must be watched, not just availability.
Security and governance implications
- Limit who can read traces and logs, because they may contain user prompts and responses.
- Assign Monitoring Contributor only to those who configure diagnostic settings.
Troubleshooting signals
- Missing traces after deployment often mean no Application Insights resource is connected to the project.
- Rising Time to Response with high utilization points to capacity, while rising response size points to prompt or output growth.
More detail
- Examine continuous monitoring in Foundry.
- Monitor latency, throughput, and response-time metrics.
- Track token consumption and resource cost.
- Configure logging, tracing, and debugging for production troubleshooting.
Ready for the quiz?
- What does connecting Application Insights to a project enable?
- Which metric best represents responsiveness?
- Which metrics reveal token-driven cost changes quickly?
- Where do diagnostic logs go, and on what scope are they configured?
Related objectives
- D4.2.S1 — Examine continuous monitoring in Foundry
- D4.2.S2 — Monitor performance metrics, including latency, throughput, and response times
- D4.2.S3 — Track and optimize cost metrics, including token consumption and resource usage
- D4.2.S4 — Configure detailed logging, tracing, and debugging capabilities for production troubleshooting