An agent returns an incomplete answer. Did a document or tool fail? Available, protected run records can help investigate the cause. Evaluate answer quality separately.
Follow one request through a run
Imagine a facilities team asks an agent: “Review this equipment replacement request against the current policy and inventory record. Flag missing details and prepare a recommendation for review. Do not place an order.” The agent may read the request, the approved replacement policy, and the inventory record. It can prepare a draft, but a person decides what happens next.
The inventory lookup times out. The trace records the failed lookup, and the draft remains incomplete.
Illustrative fields and timings, not an SDK schema or measured result. All rows: run demo-042, configuration v3. Source categories omit request contents and personal data.
Scroll horizontally to see all columns.
| Step | Source or operation | Outcome | Duration | Retries |
|---|---|---|---|---|
| 1 | Read submitted request | Success | 120 ms | 0 |
| 2 | Read approved replacement policy | Success | 80 ms | 0 |
| 3 | Read inventory | Timeout | 2,000 ms | 0 |
| 4 | Prepare review draft | Incomplete: inventory unverified | 350 ms | 0 |
Final state: The automated run stopped with an incomplete draft; the business decision remains pending review. No order operation was invoked. The trace shows that the inventory read timed out. It does not explain why or tell you whether the draft is accurate.
Inventory-timeout rate
Runs with at least one inventory timeout ÷ runs that attempted an inventory lookup in the same reporting period.
- Count each run once, even if it retries.
- Exclude runs that never attempted an inventory lookup.
- Report how many runs lack telemetry.
Illustrative day: 4 / 50 = 8%. This is an operational measure, not a measure of answer quality.
A run produces events for an operator to investigate.
Choose signals that answer operating questions
Scroll horizontally to see all columns.
| Signal | What it can help answer | What to verify |
|---|---|---|
| Run identity and configuration version | Which task and setup produced this result? | Can the team tie a run to the job, date, and active configuration without exposing unnecessary content? |
| Step and tool outcomes | Which sources were read, which tools ran, and where did the workflow stop? | Are success, failure, and skipped steps distinguishable? Are permission denials visible? |
| Timing and retries | Was the run delayed, repeated, or left incomplete? | Can the team tell a safe retry from a repeated action with side effects? |
| Handoff and completion state | Did the case go to a person or another system, and was it accepted or finished? | Do the available records distinguish “sent,” “accepted,” “completed,” and “unresolved”? Use the agent handoff guide to define the receiving system’s acknowledgment and completion states. |
| Resource and usage signals | Which runs consumed more computing resources or used more external services than expected? | What resource use is recorded, when is it available, and how does it relate to billing records? |
A trace follows the request to the timeout. A metric shows how often it happens. An event log records the failed lookup.
Check what your system captures: some omit inputs, tool results, or entire record types. Access, retention, and export settings govern what you can inspect.
Keep observability separate from evaluation
Use run records to ask what happened: which source failed, whether a tool was called, or when a handoff occurred. Use an evaluation set and reviewer criteria to ask whether the result met the job. A complete trace can still document a poor answer; a run with a short trace may still produce a useful draft.
Before a pilot, choose a few cases that matter: a routine request, an unavailable source, conflicting information, and a request outside the job. For each, name the signal that would reveal the outcome and the person who will investigate it. Review cost or usage alongside work completed and checked, not as a standalone sign of quality. See how to evaluate an AI agent and governance and cost controls.
Protect the record as well as the workflow
Run records can contain prompts, source excerpts, identifiers, tool results, or other sensitive information.
- Keep the fields needed to diagnose the run.
- Limit access and set a retention rule appropriate to the work.
- Mask or omit unnecessary sensitive data.
If records cannot explain a consequential decision, add a review step. Missing telemetry does not prove that nothing happened.
In AgentShelf
Investigate runs using available activity, error, usage, and cost records on supported paths. Provider availability, record coverage, failure handling, and retention vary by configuration.