An AI agent can produce the right answer for the wrong reasons.
It may retrieve a customer record it was not allowed to see, call the wrong tool, recover accidentally, and still end with a plausible summary. A final-answer score marks the run as successful. An operations team looking at the trace sees a near incident.
The reverse also happens. An agent may follow policy, identify that required evidence is missing, and escalate to a person. A benchmark expecting an automatic answer calls that a failure. The business should call it controlled behavior.
Agent evaluation therefore has to examine the path between goal and outcome. Observability has to preserve enough of that path in production to explain what changed, where a run failed, and whether the system remains within acceptable limits.
This is measurement for a software system whose control flow is partly probabilistic.
The short answer
Evaluate an AI agent at three levels:
Component quality: instructions, retrieval, models, tools, policies, and data. Trajectory quality: sequence of decisions, tool calls, arguments, state changes, approvals, retries, and stopping. Outcome quality: whether the task helped the user and produced the correct business state at acceptable risk, latency, and cost.
Before launch, use representative datasets, deterministic assertions, task-specific rubrics, adversarial cases, and human review. After launch, trace real sessions, monitor segmented outcomes, investigate exceptions, and turn reviewed failures into regression tests.
Evaluation, monitoring, and observability are different
These words are often used interchangeably. They solve different problems.
Evaluation
Evaluation asks whether the system meets a defined quality standard on known scenarios. It compares versions, models, prompts, policies, retrieval settings, or tools.
Monitoring
Monitoring tracks selected production measures and alerts when they cross thresholds: error rate, latency, cost, refusal rate, tool failure, or complaint volume.
Observability
Observability gives the team enough evidence to understand internal behavior from traces, logs, metrics, events, versions, and outcomes. Monitoring may tell you that completion fell. Observability helps reveal that a retrieval deployment removed tenant metadata from one document type.
Audit
Audit preserves evidence of who or what accessed data, made a decision, approved an action, and changed a system. Auditability overlaps observability but serves accountability, security, support, and compliance needs.
A dashboard without evaluation tells you the system is active, not whether it is good. An evaluation suite without production observability proves behavior in the lab, not in the conditions users create.
Why ordinary software testing is necessary but
insufficient
Traditional tests remain essential. Tool schemas, authorization, business rules, state transitions, idempotency, and deterministic code should have unit, integration, and end-to-end tests.
The incomplete belief is that those tests can prove agent behavior in the same way. A model may choose different valid paths, generate semantically equivalent text, or encounter inputs not reducible to exact strings. A brittle assertion such as "response equals this paragraph" creates false failures while missing unsupported claims with similar wording.
The recommended approach combines exact assertions with semantic rubrics:
• exact checks for permissions, tool names, fields, limits, state, and citations; • rubric grading for correctness, evidence use, completeness, tone, and usefulness; • human review for consequential, ambiguous, or disputed cases; • downstream outcome checks for what actually happened.
Strong opinion: model-based graders are useful measurement instruments, not ground truth. Calibrate them against qualified human judgments and deterministic facts.
Define the unit of evaluation
Teams often choose "conversation" because chat logs are easy to collect. The business unit may be a case, order, lead, claim, support resolution, research brief, or approval.
Define:
• start event; • goal; • initial state; • permitted data and tools; • expected stopping condition; • successful business state; • unacceptable side effects; • time and cost budget; • human handoff criteria.
This prevents a fluent conversation from hiding an incomplete task.
For long-running agents, one task may span multiple sessions and approvals. Use a durable task ID that connects every trace and outcome.
Build a failure taxonomy before choosing metrics
Goal interpretation failure
The agent pursues the wrong objective or silently changes scope.
Context failure
Required data was absent, unauthorized data was included, or stale memory distorted the task.
Retrieval failure
The system missed relevant evidence, retrieved noise, or returned a source outside the user's permissions.
Reasoning or instruction failure
The correct evidence was present, but the system reached an unsupported conclusion or ignored policy.
Tool-selection failure
The agent chose the wrong capability, called a write tool unnecessarily, or failed to use a required tool.
Argument failure
The tool was correct but inputs were invalid, incomplete, unauthorized, or commercially wrong.
Execution failure
The external action timed out, partially completed, duplicated, or returned an ambiguous state.
Escalation failure
The agent should have asked for approval or a person but continued, or escalated harmless routine work unnecessarily.
Stopping failure
The agent stopped before success, looped, continued after completion, or exceeded its budget.
Communication failure
The business state is correct but the user receives an unclear, unsupported, unsafe, or inappropriate explanation.
Outcome failure
The task appears complete but creates rework, complaint, policy breach, financial loss, or no real value.
Metrics should make these categories visible. A single accuracy number cannot.
Design an evaluation dataset that resembles production
Start with real cases
Collect representative work after appropriate privacy review. Remove unnecessary sensitive data. Preserve structure, missing fields, ambiguity, and messy language; excessive cleaning destroys realism.
Segment by risk and difficulty
Tag normal, edge, high-value, sensitive, adversarial, multilingual, and dependency-failure cases. Also tag customer type, region, tool, data source, and policy version where relevant.
Include non-answer cases
Some cases should refuse, ask a clarifying question, request approval, wait, or escalate. A dataset containing only answerable requests teaches the team to punish safe behavior.
Define expected state, not only expected text
Record the authorized records, required evidence, allowed tools, prohibited actions, expected status change, and acceptable response characteristics.
Separate development and held-out sets
Developers will learn the evaluation examples. Maintain a blind set and rotate or expand it. Use production failures to create new cases only after they are reviewed and sanitized.
Version everything
Dataset, rubric, expected outcome, policy, model, prompt, retrieval, tool schema, and grader versions should be recorded. A score without configuration lineage is not reproducible.
A layered metric framework
Task metrics
• completion rate; • correct outcome rate; • partial completion; • unnecessary escalation; • missed escalation; • human correction; • rework or reopen; • downstream success.
Retrieval metrics
• recall@k; • precision@k; • source freshness; • permission correctness; • citation support; • evidence coverage.
Tool metrics
• correct tool selection; • argument validity; • authorization compliance; • write-action approval; • execution success; • duplicate action rate; • recovery success.
Policy and safety metrics
• prohibited action attempts; • prompt-injection resistance; • sensitive-data exposure; • policy adherence; • safe refusal quality; • appropriate human handoff.
Operational metrics
• p50 and p95 task latency; • model and tool call count; • retries and loops; • tokens and cost per task; • queue time; • availability; • provider and dependency errors.
Experience metrics
• user acceptance; • edit distance or correction effort; • complaint rate; • time saved after review; • perceived usefulness; • abandonment.
Every metric needs a denominator. "Twenty tool errors" means little without total calls, task types, and affected users.
Grade the trajectory
A trajectory is the ordered sequence from goal through context, decisions, tool calls, observations, state updates, approvals, and stopping.
Trajectory evaluation asks:
Did the agent formulate the correct objective? Did it access only permitted context? Was each tool necessary and authorized? Were arguments correct at the time of execution? Did it validate tool results? Did it recover from failure safely? Did it ask for approval at the right point? Did it stop when success or a limit was reached? Did the final explanation match the actual state?
Do not reward a shorter trajectory automatically. Fewer steps may reduce cost, but skipping evidence or validation is not efficiency. Likewise, a long trace may indicate cautious investigation or a loop. Score necessity and correctness.
Deterministic checks, human graders, and model
graders
Deterministic assertions
Use code for facts: tool called, schema valid, record authorized, amount below limit, citation exists, state changed once, loop under maximum, and required field present.
These checks are cheap, repeatable, and precise. They cannot judge every semantic quality.
Human grading
Use qualified reviewers for domain correctness, nuanced policy, harmful impact, and user usefulness. Create rubrics with examples and calibration sessions. Measure agreement; people are not perfectly consistent.
Model-based grading
Use an independent or appropriately configured model to scale semantic assessment. Require structured reasons and evidence. Calibrate against human-labeled samples and monitor grader drift.
Avoid a circular setup where the same model generates, grades, and decides promotion with no external check. It may share blind spots.
Downstream verification
Whenever possible, use real outcome data: resolved ticket, accepted draft, correct ledger state, successful booking, no duplicate, or verified research claim. Outcome data can arrive late and may contain confounders, but it grounds the evaluation in work.
A realistic example: customer-support agent
An agent receives a message, retrieves account and product context, decides whether it can resolve the case, uses approved tools, drafts the response, and escalates exceptions.
Weak evaluation
A reviewer reads 50 final replies and rates 44 as good.
Stronger evaluation
For 200 segmented cases, the team measures:
• correct account and tenant context; • retrieval of the authoritative policy; • supported claims; • correct intent and issue type; • permitted tool selection; • valid arguments and idempotency; • approval for credit or account change; • escalation for complaints, safety, and missing evidence; • final system state; • response usefulness; • latency and cost; • ticket reopen within seven days.
One case produces an excellent reply but used another customer's document through a faulty cache. Final-answer grading says pass. Permission evaluation says critical fail.
Another case refuses to process a refund because the account record conflicts with the request and routes it to billing. A naive completion metric says fail. The policy rubric says pass.
This is why the metric hierarchy matters.
Production observability architecture
Correlation and identity
Assign request, session, task, user, tenant, and operation identifiers. Preserve identity without exposing unnecessary personal data in logs.
Version context
Record model, prompt, policy, retrieval, tool, data-source, and grader versions. Store feature flags and experiment cohort.
Trace spans
Represent retrieval, model call, tool call, approval, retry, and external dependency as timed spans. Include structured outcomes and error types.
Evidence references
Store source identifiers, access decision, and freshness rather than copying every sensitive document into telemetry.
Business outcome
Connect technical traces to ticket, order, case, or workflow outcome. Without this connection, teams optimize response metrics rather than business value.
Privacy and retention
Redact or tokenize sensitive fields, restrict trace access, set retention, support deletion, and separate security audit from product analytics where appropriate.
Alerts that deserve attention
Useful alerts are segmented and actionable:
• sudden increase in tool denial; • cross-tenant or unauthorized retrieval attempt; • completion drop after a model or prompt release; • rising correction rate for one policy; • cost per task above budget; • loop or retry spike; • p95 latency breach; • approval queue aging; • unsupported claims in high-risk segment; • unexpected new tool sequence; • user complaint correlated with a version.
Avoid alerting on every low-quality output. That creates noise and trains operators to ignore the system. Define severity, owner, response, and service level.
Regression and release gates
Run the evaluation suite when any material component changes:
• model or provider; • system prompt or policy; • tool description or schema; • retrieval model, chunking, index, or reranker; • source data or permission logic; • orchestration or retry policy; • approval rule; • user interface that changes reviewer behavior.
Compare candidate and production versions on the same held-out set. Require no critical regression even if the average score improves. Use canary release and watch production segments.
One warning: prompt optimization against a static benchmark can overfit just like model training. Maintain blind cases and periodically refresh the suite.
Failure review and continuous improvement
For each significant production failure:
Contain the impact. Preserve trace and configuration evidence. Identify the failure layer. Check whether the evaluation suite contained a similar case. Correct source, policy, prompt, tool, model, or process. Add a reviewed regression case. Test for side effects. Release through the normal gate. Monitor the affected segment.
Do not automatically retrain on every failure. Some failures come from bad data, ambiguous policy, outages, or a misleading interface. Learning the wrong lesson is still regression.
Common mistakes
Scoring only the final response
This misses unauthorized access, wrong tools, duplicated actions, and lucky recovery.
Using one average score
Critical failures disappear inside high-volume easy cases. Segment by consequence and workflow.
Treating model confidence as quality
Confidence is not evidence of correctness or business safety.
Using synthetic examples only
Synthetic data helps coverage but rarely reproduces operational mess, missing context, and organizational edge cases on its own.
Trusting one grader
Calibrate automated graders and sample human review. Use deterministic facts whenever possible.
Logging everything forever
Observability can become a sensitive data leak. Minimize, protect, and expire.
Improving the model when the tool is broken
Diagnose the layer before changing prompts or models.
Ignoring human-review quality
Reviewers need context, calibration, capacity, and a way to disagree. Human labels are not automatically correct.
A practical implementation framework: TRACE
T - Task and target state
Define the business unit, success state, prohibited side effects, risk segments, and owner.
R - Representative cases
Collect real, edge, adversarial, failure, escalation, and non-answer scenarios. Version the set.
A - Assertions and annotations
Combine deterministic checks, task-specific rubrics, human calibration, model graders, and downstream verification.
C - Correlated production evidence
Trace users, tenants, versions, retrieval, tools, approvals, latency, cost, and outcomes with privacy controls.
E - Evolution gates
Turn reviewed failures into regression tests, compare versions, canary releases, and monitor affected segments.
Establish a baseline before the agent
Agent metrics are misleading without a comparison point. Measure the current human or rules-based process on the same business outcome: correct completion, cycle time, rework, escalation, cost, and consequence of errors. The baseline will not be perfect. It still prevents the team from celebrating a 90% agent score when the existing process achieves 97% on the cases that matter.
Compare like with like. Humans may handle only difficult cases after automation removes easy ones, making their post-launch accuracy appear worse. Agents may receive cleaner inputs during a pilot than production staff receive. Segment by case type, channel, risk, and data quality before drawing conclusions.
A useful rollout measures three populations for a limited period: the existing process, the agent in shadow or recommendation mode, and the final human-plus-agent process. This exposes whether value comes from model decisions, faster retrieval, better routing, or simply a redesigned workflow.
Do not treat the human baseline as the ceiling. It is a reference. The target should reflect business risk and what the combined system can reliably achieve.
Recalculate that baseline whenever process rules, staffing, or input quality materially change.
Frequently asked questions
What is AI agent evaluation?
It is the structured measurement of whether an agent completes tasks correctly and safely across its components, trajectory, and business outcome.
How is agent evaluation different from LLM evaluation?
LLM evaluation may score a single model response. Agent evaluation also measures retrieval, tool selection, arguments, state, approvals, retries, stopping, and external actions.
What is AI agent observability?
It is the ability to understand agent behavior in production through correlated traces, metrics, logs, events, versions, evidence, and outcomes.
What metrics should be used for AI agents?
Use task success, retrieval quality, tool correctness, policy, escalation, stopping, correction, downstream outcome, latency, cost, and incident measures. Segment by risk.
What is trace grading?
Trace grading evaluates the sequence of steps an agent took, not only the final output. It identifies where retrieval, reasoning, tool use, or escalation failed.
Can another AI model grade an agent?
Yes, for scalable semantic assessment, but calibrate it against human and deterministic evidence. A model grader is an instrument, not unquestioned truth.
How many evaluation cases are needed?
There is no universal number. Begin with enough representative cases to cover major paths and risks, then expand from reviewed production failures. Coverage and quality matter more than raw count.
How often should agent evaluations run?
Run them for material changes and on a regular regression schedule. Monitor production continuously and investigate segmented shifts.
Should production conversations become training data?
Not automatically. Review privacy, consent, quality, policy, and provenance. Sanitize and approve examples before using them for evaluation or training.
What is the most important agent metric?
The correct business outcome without unacceptable side effects. No single metric captures that, so use a hierarchy of task, risk, operational, and experience measures.
Conclusion
An AI agent is not a text generator with a longer prompt. It is a sequence of decisions and actions inside a business system.
Evaluation must therefore inspect components, trajectory, and outcome. Observability must connect production behavior to versions, evidence, tools, people, cost, and final state. The goal is not to collect the most telemetry. It is to make failures diagnosable and improvements defensible.
Measure what the agent did. Measure what changed in the business. Keep the two connected.
Actionable next steps
Define the business task and successful final state. Create a failure taxonomy for the workflow. Collect 50 to 200 representative, segmented cases. Add deterministic assertions for permissions, tools, state, and limits. Calibrate semantic rubrics against qualified human review. Trace the full trajectory with version and outcome context.
Build release gates that block critical regressions. Turn reviewed production failures into new tests.
Wizora Studio's AI agent development work includes task definition, tool boundaries, evaluation scenarios, activity traces, approvals, and monitored rollout. A useful assessment begins with actual cases and the consequences of a wrong action, not a generic benchmark.
References
• OpenAI - AgentKit and Trace Grading • OpenAI - A Practical Guide to Building AI Agents • NIST - AI Metrology Center • AWS - Generative AI Lens: Operational Excellence • AWS - Operational Excellence: Prepare
Editorial note: the support scenario is a composite example and does not claim Wizora Studio client results.
AI Agent Evaluation and Observability: Measure the Work, Not Just the Answer: a practical decision framework
AI Agent Evaluation and Observability: Measure the Work, Not Just the Answer should be evaluated against the real problem, the intended audience, the systems involved, and the level of human review required. The right approach is the one that makes the workflow more useful and more inspectable, not the one that simply adds another tool or trend to the stack.
Key topics to cover
- AI Agent Evaluation and Observability: Measure the Work, Not Just the Answer guide
- AI Agent Evaluation and Observability: Measure the Work, Not Just the Answer best practices
- AI Agent Evaluation and Observability: Measure the Work, Not Just the Answer examples
- AI Agent Evaluation and Observability: Measure the Work, Not Just the Answer architecture
Use these topics as supporting language only when they answer a real question in the article. Explain the implementation choices in plain language, distinguish a reliable workflow from a prototype, and qualify claims that depend on the project scope, data quality, vendor limits, or operating model.
Questions readers should ask
- Why ordinary software testing is necessary but
- The short answer
- How does AI Agent Evaluation and Observability measure the work?
Limits, evidence, and next steps
Results depend on the workflow, inputs, integrations, security requirements, and review process. Do not treat this guide as a guarantee of cost, speed, rankings, compliance, or business outcomes. Document the assumptions, define what will be measured, and keep a clear stopping or escalation condition.
If you want to map the topic to a real project, review the relevant Wizora service or contact Wizora Studio with the current process, constraints, and desired outcome.



