Measure a website agent on whether its answers help visitors and stay within approved sources, then assess whether handoffs reach the right team with enough context. Keep these outcomes separate from conversation volume or length; more interaction does not by itself mean better service.
Fictional example: An agent for Harborview Workshop helps visitors ask about a training room. Its public pages state a seated capacity of 36, describe a step-free entrance, and explain how to request a date. The agent can answer those facts and prepare an inquiry, but cannot check the live room calendar or reserve the space.
Use separate cases to inspect answer quality, boundaries, and recovery.
Choose measures that show whether visitors got help
Start with the result the visitor needs. For Harborview, choose measures such as:
- Supported answer rate: What share of sampled factual replies has every material claim supported by the current approved pages?
- Answer usefulness: Does the response answer the question, use the right location or service detail, and state important limits?
- Correct next step: When availability is unknown, does the agent offer the request path without implying a reservation?
- Handoff completeness: Does the team receive the visitor’s question, relevant source, missing decision, and chosen contact details?
- Handoff acceptance: What share of submitted handoffs receives an acknowledgement of ownership within the agreed response window?
- Follow-up completion: What share of accepted handoffs receives the promised follow-up by its deadline?
Define each measure with a numerator, denominator, and review method. Supported answer rate is the number of sampled factual answers with every material claim supported, divided by all sampled factual answers. A reviewer decides which claims matter and what counts as support; an answer with one unsupported material claim does not pass.
For handoffs, define a submission cohort and response windows before measuring. For example, allow one business day to acknowledge ownership and two business days after acceptance to complete follow-up. Calculate the rates only once every case in the cohort has had its full applicable window; show younger cases as pending.
An illustrative weekly sample—not a product benchmark—could contain 20 factual answers, with 18 fully supported: 18 / 20 = 90%. Of 10 submitted handoffs, 8 are accepted within one business day: 8 / 10 = 80%. If 6 of those 8 receive follow-up within two business days of acceptance, completion is 6 / 8 = 75%. The other two accepted cases remain overdue or unresolved; acknowledgement alone does not count as follow-up.
Score a concrete example
Suppose a visitor asks, “Can 36 people sit in the room, is the entrance step-free, and can I book October 16?” The agent checks the capacity, accessibility, and request pages. A suitable reply is:
“The room page lists seating for 36, and the accessibility page describes a step-free entrance. I can prepare a request for October 16, but the room is not reserved until the team checks the calendar and confirms it.”
The answer is useful if the cited pages support the capacity and entrance details and the visitor can reach the request form. Check packet completeness when the team receives the date, relevant question, and chosen contact details. Record acceptance when the team acknowledges ownership, and follow-up completion when it responds to the visitor. None of those events proves a booking took place.
The following single-case scores are illustrative, not a product benchmark. If the visitor asks the team to check the date, score the answer and handoff separately:
Scroll horizontally to see all columns.
| Outcome | Check for this visitor | Illustrative score |
|---|---|---|
| Usefulness | The reply addresses capacity and step-free access, offers a way to request October 16, and says no booking is confirmed. | 3 / 3 checks |
| Approved support | The capacity and entrance claims match their current approved pages; availability remains unconfirmed. | 2 / 2 claims supported |
| Handoff quality | The draft records the October 16 question, pages checked, unresolved calendar decision, and room team as owner. The visitor has not chosen a follow-up route. | 4 / 5 fields; keep the draft pending |
Test common and difficult cases
Build a small evaluation set from visitor needs: a direct question, a paraphrase, an ambiguous request, an answer absent from the site, conflicting page details, a request to guarantee a date, and a case that needs a person. Record the expected answer, source, next step, and failure behavior for each. Test response quality and routing separately so a strong answer cannot hide a broken handoff.
During a pilot, sample real interactions only under the organization’s privacy and retention rules. Limit what is collected, restrict who can inspect it, and redact details that reviewers do not need. Ask people who own the visitor task to review a consistent sample; automated scores can help organize review but do not decide whether the service outcome is acceptable.
Choose a decision rule before expanding
Set release thresholds with the people responsible for the service, review capacity, and consequences of a wrong answer. Define which errors stop the pilot, which require a corrective change, and who approves expansion. If a source, prompt, or handoff route changes, retest the cases affected by that change. Report the sample and its limits alongside the result rather than presenting one score as proof of overall quality.
Start with agent evaluation, define the human handoff, and review the website agent pattern before widening the visitor experience.