Guide

AI Agent Failure Modes and Recovery

Diagnose an agent failure from evidence, then choose a bounded retry, correction, safe stop, or human handoff.

By AgentShelfUpdated September 28, 2026

Agentshelf Knowledge Guide

An agent failure can come from missing information, an unavailable tool, conflicting sources, an unsupported action, or a result that does not meet the job. Diagnose the cause from the run evidence before choosing a response; repeating the whole task is not a safe default.

Start with one equipment request

A facilities coordinator asks: “Check this equipment replacement request against the current policy and inventory record. If required information is missing, draft a question for review. Do not place an order.” The agent can read the submitted request, approved policy, and inventory record. Its only output is a recommendation draft.

If the inventory system is temporarily unavailable, the agent cannot confirm whether the item is in stock. A responsible result would be:

Replacement review — incomplete draft

  • Checked: The submitted request and current replacement policy.
  • Could not verify: Current inventory; the source did not respond.
  • Decision: Do not recommend a replacement until inventory is checked.
  • Next step: Retry the read later if the system confirms it is safe, or route the case to facilities.

The draft keeps the verified work useful while making its limit visible. It does not claim that the inventory lookup succeeded or that a replacement is approved.

Illustrative evaluation scorecard separating outcome, process, boundaries, and recovery for an agent workflow

Record what failed before deciding how to recover.

Diagnose the failure before acting

Scroll horizontally to see all columns.

SymptomEvidence to checkResponse to consider
A source is missing, unreadable, or staleSource identifier, version or date, retrieval result, and any access errorMark the result incomplete. Ask the source owner to restore or update the reference, or hand off the case.
Two approved sources disagreeThe exact fields and dates that conflict, plus the source owner for eachShow the disagreement and pause the decision until an authorized owner resolves it. Do not choose whichever value looks more plausible.
A tool times out or returns an errorTool name, operation, error category, retry history, and whether an earlier attempt changed anythingRetry only when the failure appears temporary and repeating the operation is safe. Use a bounded retry policy; otherwise stop and route it.
The output is incomplete or unsupportedRequired fields, source evidence, and the evaluation case that detected the gapCorrect the draft or ask for review. Keep an unsupported result from flowing into a later action.
A request or document tries to change the jobThe original task, untrusted text, selected tool, and permission decisionTreat the text as evidence to assess, not as authorization. Deny or pause any action outside the defined job.

Use these questions to investigate the failure. The records and error details available will vary by system, so check what your run history retains and what the receiving tool reports.

Choose a recovery path with a stop condition

  1. Preserve the evidence. Keep the request, active configuration, source versions, tool outcome, and visible error needed to explain the failure. Protect sensitive records under the team’s access and retention rules.
  2. Classify the condition. Separate a temporary dependency problem from a missing reference, policy conflict, invalid input, or unsupported request. Ask the relevant source or system owner when the cause is uncertain.
  3. Choose the smallest safe response. Retry a read only when it is safe and bounded. Do not repeat a write or external action until its prior status is known. Correct a draft only after the evidence is available; otherwise stop and hand off.
  4. Make the state visible. Tell the requester what was checked, what remains unknown, and who owns the next step. Do not label a partial run as complete.

For the equipment case, a failed read should leave the recommendation pending. A facilities reviewer can check the inventory record, then either continue or close the request. If the system cannot tell whether an action already ran, the operator should resolve that state before trying again.

Retest the failure path

Keep the original case that exposed the problem and add a nearby variation: an unreadable source, a delayed response, a conflicting record, and an out-of-scope request. Confirm both the final result and the path taken. A retry that produces an answer may still be wrong if it repeats a side effect, uses an unapproved source, or hides the uncertainty.

Review the evidence with agent evaluation, observability for agent runs, and the security review checklist.

Your privacy choices

We use optional assistant personalization, analytics, and advertising technologies only when you allow them. Necessary site functions remain active. Cookie Policy