An agent evaluation checks whether the configured workflow does its job, uses only allowed information and actions, and handles failure safely. For example, imagine a Northstar Supplies onboarding agent asked to prepare a procurement brief from an intake packet, an approved checklist, and the current vendor record. It must flag a missing insurance certificate and draft a request for review; it must never approve the vendor or send the request on its own.
Illustrative scorecard, not measured results.
In this sample:
- Routine packet: Passes all four checks.
- Missing evidence: Produces a useful draft within its limits, but its process and recovery need review.
- Conflicting records: Fails because the draft neither resolves nor clearly hands off the disagreement, even though it respects permissions.
- Tool failure: Cannot complete the requested brief, so its outcome fails. Its process, limits, and recovery pass because it reports the failed lookup and hands off safely.
A safe handoff and a completed task are separate results.
For this illustrative release decision:
- Stop: A permission breach or unknown side-effect state blocks the run; resolve it before resuming.
- Revise: A correctable evidence or handoff gap must be fixed and retested before widening.
- Widen: Consider it only after affected cases meet their written expectations and the owner accepts the evidence.
Decision for these sample results: revise before widening. The conflicting-record case does not hand off clearly, and the missing-evidence case needs procedure and recovery review. The tool-failure case passes process, boundaries, and recovery while its business outcome remains incomplete.
Define the expected result
Before testing, write down what a good run must contain and what it must not do. For this example, the agent should compare the packet with the checklist, ground each finding in the supplied materials, call out the missing certificate, and leave both the approval and outgoing message to a person. A polished brief that invents evidence or sends the request is a failed run even if its wording looks convincing.
Choose a workflow owner and a reviewer who can judge whether the brief is useful. Agree which errors block a pilot, which can be handled with review, and which can be tracked. Set thresholds for the consequences and the team’s review capacity.
Write these expectations into a focused agent job brief so the task, evidence, allowed actions, and stop conditions stay fixed while you test.
Test a representative set
Use a small set of safe, representative packets:
- One complete packet.
- One missing the certificate.
- One with conflicting details.
- One with an unreadable attachment.
Record the expected result and why it is correct for each packet. Include varied phrasing and requests outside the agent’s job. Use synthetic or appropriately protected data in the test environment.
Keep the task and success criteria fixed as the packet changes. This makes it easier to see whether a result changed because of the evidence, the prompt, or the configured tools. Repeat runs when the same input can produce different results, and inspect the steps as well as the final brief.
Inspect the run and draft
Check which sources the agent read, what it extracted, whether it used the permitted lookup, and whether it stopped before approval or sending. Review the returned draft for missing facts, unsupported claims, and clear ownership of the next step.
Illustrative output using fictional Northstar Supplies records:
Northstar Supplies — onboarding draft
- Finding: A current insurance certificate is missing from the intake packet.
- Draft request: “Please provide a current insurance certificate so procurement can continue its review.”
- Decision: Vendor approval has not been made; a reviewer must approve any outgoing request.
This sample is not a real vendor record or a measured system result. In a test, verify that the missing-document finding matches the packet and checklist, that the message stays a draft, and that no approval or send action occurred.
Probe failures and score the result
Test a missing source, an unreadable file, a tool error, conflicting records, an irrelevant instruction inside a document, and an unanswered or rejected approval. The agent should explain what it could not verify and route the case as designed.
Score four things separately:
Scroll horizontally to see all columns.
| Measure | What to check |
|---|---|
| Outcome | The brief’s accuracy and usefulness. |
| Process | Whether the agent followed the allowed sequence. |
| Boundaries | Whether it stayed within its permissions. |
| Recovery | Whether it recovered or handed off clearly. |
Use direct checks for required fields and unchanged records, and human review for judgments that need subject knowledge. Compare automated grades with reviewer decisions before using them to triage runs.
Retest before widening access
Keep the task set, agent version, configuration, results, reviewers, limitations, and decision together. When something changes, rerun the case that exposed the issue and nearby cases. Treat results as evidence about the tested setup and scenarios, then identify what still needs human review before expanding the pilot.
For related checks, read AI agent governance and cost controls and permissions and boundaries.