How we test an AI agent before it talks to your customers

Real past cases, agreed pass marks and a test run after every change. A look at the testing process behind every agent we launch.

Acropine team7 min read

A traditional program gives the same output for the same input every time. An AI agent doesn’t, quite. That makes testing more important, not less. Here is the process we follow for every agent we launch.

1. Build a test set from real cases

We start by collecting a sample of real, past cases from your business: support tickets, call recordings, invoices, WhatsApp enquiries. We deliberately include the awkward ones: angry customers, poor scans, mixed languages, questions with no good answer.

For each case, your team tells us what a good outcome looks like. That becomes the answer key.

2. Agree the pass marks

Before we build anything, we agree with you what the agent has to achieve. For example:

  • extracts invoice totals correctly in at least 98% of cases
  • never promises a refund without approval
  • hands over to a person when a caller asks for one, every time

Some rules are about accuracy. Others are hard limits that must never be broken.

3. Run every case, every time

Each time we change the agent, whether that’s a new prompt, a new model or a new integration, we run the whole test set again automatically. We compare the results with the previous version, so an improvement in one place can’t quietly break something elsewhere.

4. Review the failures by hand

Scores only tell part of the story. We read through every failed case to understand why it failed. Sometimes the agent is wrong. Sometimes the answer key is, because a policy changed. Both get fixed.

5. Launch in stages

We don’t switch everything on at once. A typical launch starts with the agent drafting answers for your team to approve, then handling a small share of real traffic on its own, then more as the numbers hold up.

6. Keep adding to the test set

After launch, every mistake the agent makes in real use is added to the test set. Over time, the test set becomes a detailed record of your business’s edge cases, and the agent is checked against all of them before any change goes live.

This is slower than shipping a demo. It is also why the agents we launch stay live.