Prepared with AI assistance. Examples and checklists are illustrative guidance, not hands-on test results or hiring guarantees. The accompanying image is an AI-generated editorial illustration.
Quick answer: Start with one support task, a small set of realistic cases and explicit failure rules. Evaluate correctness, evidence, escalation and review effort before allowing customer-facing use. A polished demonstration is not enough to establish reliability.
An assistant that drafts a pleasant reply can still invent a refund policy or overlook an account-security concern. A useful pilot asks whether the system behaves appropriately when the answer is missing, ambiguous or outside its authority.
NIST’s generative AI risk guidance identifies risks including confidently incorrect content. The testing routine below is a practical starting template for a small team; it is not a certification or a substitute for the controls required in your organisation.
Choose one narrow job
A sensible first task is drafting answers to a clearly documented, low-risk product question for an agent to review. Define what the assistant can read and what it must not do. For the initial pilot, do not give it authority to issue refunds, change accounts or send messages automatically.
The distinction matters: the NCSC notes that agentic systems can use tools and take actions. Adding that ability changes the risk. A draft-only pilot tells you little about whether an autonomous workflow is safe.
Build a starter set of 20 cases
The number 20 is an organisational suggestion, not a statistically sufficient sample or industry benchmark. Start small enough to inspect every answer, then expand based on the real variety of your support queue.
| Case type | Starter count | What to test |
|---|---|---|
| Routine documented questions | 6 | Correct answer grounded in the current help material |
| Incomplete questions | 4 | Useful clarification instead of guessing |
| Conflicting or outdated information | 3 | Recognition that sources disagree |
| Requests outside policy | 3 | No invented exceptions or promises |
| Sensitive or account-specific requests | 2 | Escalation to the approved process |
| Instructions trying to override the task | 2 | No compliance with untrusted instructions inside customer text |
Use synthetic cases or appropriately approved and de-identified material. Preserve realistic ambiguity without uploading confidential tickets into an unapproved service.
Write the expected behaviour first
For each case, record the approved source, the correct action, unacceptable claims and the escalation owner. A complete model answer is optional; the required behaviour is not.
Illustrative case: A customer asks for a refund after the published window. The expected behaviour is to explain the documented route and refer exceptional requests to the authorised team. An unacceptable response promises approval or invents an extended deadline.
This prevents a common evaluation mistake: deciding an answer is good simply because it sounds plausible after you read it.
Review four dimensions separately
- Correctness: Is the answer supported by the approved material?
- Boundary handling: Does the assistant recognise missing information and prohibited actions?
- Customer usability: Is the next step clear and appropriate?
- Review effort: How much agent work is needed to make the draft usable?
Use simple labels such as acceptable, needs correction and unacceptable. Keep privacy leaks or unauthorised commitments visible as critical failures instead of averaging them away inside a single score.
Measure the work, not just generation speed
Time the complete task: retrieval, drafting, checking, correction and recording the response. Compare with a manual baseline using similar cases. Keep the test conditions consistent and record the number of cases. A tiny internal sample can guide your next experiment, but it cannot justify a universal productivity claim.
Where possible, ask another knowledgeable reviewer to inspect the difficult cases. Discuss disagreements and improve the expected-behaviour notes before treating the result as reliable.
Define the launch decision before seeing results
Set a written rule appropriate to the task. For example, any exposed sensitive information or invented financial commitment pauses the pilot; routine errors require source or prompt changes and retesting. This is a proposed internal decision rule, not a promise that zero observed failures means zero risk.
If results justify a limited rollout, retain human approval and a clear stop control. Log corrections, repeat failed cases after changes and add new examples from real use. Keep a separate set of unseen cases so you are not only testing answers the team has already tuned.
Your next step
Write five difficult cases before choosing a vendor. A tool that handles your real boundaries is more valuable than one that performs perfectly on its own demonstration questions.
Sources and further reading
Sources checked during preparation on 26 September 2026. Product features and policies can change; consult the current source for your situation.
