An agent harness is the setup that defines a model’s job, context, available tools, limits, and review rules. It turns a broad instruction into a repeatable workflow that a team can inspect and test.
Consider a hypothetical Northstar Supplies onboarding agent. Procurement asks: “Prepare an onboarding review from the intake packet, approved checklist, and current vendor record. Flag missing evidence and draft a document request for review. Do not approve the vendor or send a message.” The harness describes what the agent may read, what draft it should produce, and where a person must take over.
A model enclosed by task controls.
Define the parts around the task
- Job: Prepare an onboarding review and draft a request for missing evidence.
- Context: Use the submitted packet, approved checklist, and current vendor record. Identify which source supports each finding.
- Tools: In this hypothetical setup, allow reading those sources and creating a draft. Do not provide approve or send actions.
- Rules: Flag missing or conflicting information, do not fill gaps by guessing, and stop before approval or sending.
- Review: Give the draft to procurement, which decides whether to request another document or continue the review.
Illustrative output from fictional Northstar Supplies records:
Onboarding review — draft
- Finding: A current insurance certificate is missing from the intake packet.
- Draft request: “Please provide a current insurance certificate so procurement can continue its review.”
- State: No message sent; no vendor approval made.
This sample shows how job, context, tools, and limits shape the result. If the checklist or current record cannot be read, the agent should report that gap and hand off rather than present an unverified conclusion. Instructions alone do not guarantee correct behavior; test the workflow against complete, incomplete, and conflicting packets.
A harness is more than a prompt
A prompt can describe the role and operating rules, but the overall setup also includes the information and actions actually available to the model, plus the environment and review process.
Control rule: A written rule that says “do not approve” is not a substitute for withholding approval capability where the system allows that control.
Without those controls, the Northstar workflow has visible failure modes. If the checklist cannot be read but the draft still says the certificate is missing, the reviewer has no checklist evidence for that finding. If the agent still has access to a send action despite its draft-only instructions, check the outgoing-message status for any request sent before review. Require evidence from the named sources, stop and hand off when a source is unavailable, and withhold approval and send actions from this workflow.
Before testing, make sure a teammate can:
- State the job in one sentence.
- Name every source.
- Explain each permitted action.
- Identify the human decision point.
Narrow the task if any of these remain unclear.
For a component map, read AI agent architecture. To define the job, see How to define a focused agent job. For current product-specific terms, consult the AgentShelf agent harness page.