Guide

AI Agent Observability: What to Track

Track workflow activity, errors, handoffs, and resource signals so an owner can investigate a run and decide what needs review.

By AgentShelfUpdated September 29, 2026

Agentshelf Knowledge Guide

An agent returns an incomplete answer. Did a document or tool fail? Available, protected run records can help investigate the cause. Evaluate answer quality separately.

Follow one request through a run

Imagine a facilities team asks an agent: “Review this equipment replacement request against the current policy and inventory record. Flag missing details and prepare a recommendation for review. Do not place an order.” The agent may read the request, the approved replacement policy, and the inventory record. It can prepare a draft, but a person decides what happens next.

The inventory lookup times out. The trace records the failed lookup, and the draft remains incomplete.

Illustrative fields and timings, not an SDK schema or measured result. All rows: run demo-042, configuration v3. Source categories omit request contents and personal data.

Scroll horizontally to see all columns.

StepSource or operationOutcomeDurationRetries
1Read submitted requestSuccess120 ms0
2Read approved replacement policySuccess80 ms0
3Read inventoryTimeout2,000 ms0
4Prepare review draftIncomplete: inventory unverified350 ms0

Final state: The automated run stopped with an incomplete draft; the business decision remains pending review. No order operation was invoked. The trace shows that the inventory read timed out. It does not explain why or tell you whether the draft is accurate.

Inventory-timeout rate

Runs with at least one inventory timeout ÷ runs that attempted an inventory lookup in the same reporting period.

  • Count each run once, even if it retries.
  • Exclude runs that never attempted an inventory lookup.
  • Report how many runs lack telemetry.

Illustrative day: 4 / 50 = 8%. This is an operational measure, not a measure of answer quality.

Illustration of an agent run producing events that an operator examines to understand what happened

A run produces events for an operator to investigate.

Choose signals that answer operating questions

Scroll horizontally to see all columns.

SignalWhat it can help answerWhat to verify
Run identity and configuration versionWhich task and setup produced this result?Can the team tie a run to the job, date, and active configuration without exposing unnecessary content?
Step and tool outcomesWhich sources were read, which tools ran, and where did the workflow stop?Are success, failure, and skipped steps distinguishable? Are permission denials visible?
Timing and retriesWas the run delayed, repeated, or left incomplete?Can the team tell a safe retry from a repeated action with side effects?
Handoff and completion stateDid the case go to a person or another system, and was it accepted or finished?Do the available records distinguish “sent,” “accepted,” “completed,” and “unresolved”? Use the agent handoff guide to define the receiving system’s acknowledgment and completion states.
Resource and usage signalsWhich runs consumed more computing resources or used more external services than expected?What resource use is recorded, when is it available, and how does it relate to billing records?

A trace follows the request to the timeout. A metric shows how often it happens. An event log records the failed lookup.

Check what your system captures: some omit inputs, tool results, or entire record types. Access, retention, and export settings govern what you can inspect.

Keep observability separate from evaluation

Use run records to ask what happened: which source failed, whether a tool was called, or when a handoff occurred. Use an evaluation set and reviewer criteria to ask whether the result met the job. A complete trace can still document a poor answer; a run with a short trace may still produce a useful draft.

Before a pilot, choose a few cases that matter: a routine request, an unavailable source, conflicting information, and a request outside the job. For each, name the signal that would reveal the outcome and the person who will investigate it. Review cost or usage alongside work completed and checked, not as a standalone sign of quality. See how to evaluate an AI agent and governance and cost controls.

Protect the record as well as the workflow

Run records can contain prompts, source excerpts, identifiers, tool results, or other sensitive information.

  • Keep the fields needed to diagnose the run.
  • Limit access and set a retention rule appropriate to the work.
  • Mask or omit unnecessary sensitive data.

If records cannot explain a consequential decision, add a review step. Missing telemetry does not prove that nothing happened.

In AgentShelf

Investigate runs using available activity, error, usage, and cost records on supported paths. Provider availability, record coverage, failure handling, and retention vary by configuration.

Your privacy choices

We use optional assistant personalization, analytics, and advertising technologies only when you allow them. Necessary site functions remain active. Cookie Policy