Published October 9, 2026 in For program leaders
How we test the AI coach before a change reaches your entrepreneurs
We ran the coach with live cohorts for months before we had a proper way to test it. Here is what stood in for one, why it stopped being enough, and what we built instead.
On this page
- What makes an AI business coach hard to test
- How we checked it at first
- Where it stopped being enough
- We wrote the checklist before the code
- What a test looks like in practice
- Trusting a second AI to grade the first
- What changed for us
- What it does not tell us
- Questions to ask anyone selling you an AI coach
- Sources
When a small business program takes on a new coach, the director has ways of knowing whether the coaching is any good. You sit in on a session. You read the notes. You ask the cohort at the end of the week how it went. None of it is scientific, and all of it works, because a coach is one person whose habits you get to know.
When the coach is an AI agent, the same questions get harder, because the AI keeps changing, for three reasons. The first is us: we edit its instructions whenever a business owner or a coach tells us something could be better.
The second is the AI underneath. The companies that make these models retire old ones and release new ones several times a year, with 60 days' notice as the minimum, so a coach built on one model has to move to the next whether or not we planned to. Even a model that stays put can shift: researchers who checked one widely used model three months apart found its accuracy on a simple maths question had dropped from 84% to 51% on the same set of questions.
The third is that a newer model is not automatically a better coach. It reads the same instructions differently, and a model maker's own migration notes list the places where wording that worked before has to be rewritten.
So a reply that was good on Monday can be worse on Thursday, and nobody sat in on Thursday's session.
This post is how we answer that question for the Proactivly AI coach: whether the coaching is any good, this week. We ran it with live cohorts for months before we had a proper way to test it. Here is what stood in for one, why it stopped being enough, and what we built instead. It is written for the director who wants to know what "we test it" means when a vendor says it, including when the vendor is us.
What makes an AI business coach hard to test
Proactivly is the AI coaching platform for small business programs. Our first deployment is with a small business training organisation in Toronto. Each cohort has 12 to 15 entrepreneurs and one human coach, who works by asking questions.
Most advice on testing AI agents assumes the agent finishes a task you can check: a ticket closed, a booking made, code that passes its tests. Coaching has no finish line like that. An AI agent holding a business coaching conversation has no right answer to compare against. A session runs over many turns, often by voice, and the same question from two entrepreneurs deserves two different replies.
The harder part is that a reply can be accurate and still be bad business coaching. If an entrepreneur asks what to charge and the AI coach names a figure, the figure may be reasonable and the reply has still failed, because the entrepreneur should have reached it through their own costs and their own customers. The coach's job is to ask the question a good coach would ask. That is what we have to test, and it is not something a script can tick.
How we checked it at first
For the first months, three things each did part of the job.
- A person played the entrepreneur. A tester sat with the AI coach for an hour at a time, pretended to be starting a bakery or a cleaning company, and wrote up what the coach got wrong in a structured form: what was asked, what came back, and what should have come back.
- Automated checks that the app still works. Scripts open the app, sign in, start a session, and confirm that the pages load, the session saves, and the coach is running on the AI model we meant it to. They run on every change, before it goes anywhere.
- A record of every call to the AI. Each reply is logged with which model produced it, how long it was, and what it cost, so a change that quietly doubled the bill shows up the same day.
Put together, this meant that every change started with a person reading replies, and ended with scripts confirming the app had not broken. The scripts could tell us a change broke something. They could not tell us it made the coaching worse. Only the tester could, and the tester could only read so many sessions.
Where it stopped being enough
As the coach grew, three things stayed unknown after every change.
- Whether the change helped. We had no score to compare before and after an edit to the coach's instructions, or a move to a newer AI model. We had a tester's impression of the sessions they happened to run.
- What we had not looked at. Nobody replayed the same conversations after each change. A session that had gone well in August was not checked again in September.
- Whether a second reader would agree. With one tester, we could not tell a bad reply from a disputed one. Some of the notes were about replies we thought were right.
The coach works from a long set of written instructions, and that set grew with every complaint. A line added to fix one complaint could quietly undo the line above it, and nothing in our setup would notice until a coach in the program did.
We wrote the checklist before the code
The temptation was to start building a testing tool. We started with a checklist instead, because the tool only grades against whatever you give it.
Every complaint a tester, a business owner or a coach had raised became a question with a yes or no answer. We tried five-point scales first and dropped them. A scale lets two readers give a 3 and a 4 to the same reply and feel they agree. A yes or no forces the question to be written precisely enough that they would answer it the same way.
The five questions we ask of every reply
- Is it about this entrepreneur's business? The reply refers to their idea, their customers and their numbers, not to a business in general.
- Does it ask before it tells? The coach draws the answer out with a question before offering one.
- Is the tone right? Encouraging, with no flattery and no pressure.
- Is it careful with claims? No financial or legal statement is presented as certain.
- Does it stay on the program's path? The reply works inside the program's curriculum rather than wandering off it.
None of the five is an engineering measure, and none started with us deciding what good coaching ought to be.
The rule for the list is short. If two readers would answer a question differently about the same reply, the question is wrong, not the readers, and it gets rewritten until they would not.
What a test looks like in practice
A test starts with a made-up entrepreneur. Each one is a short description: the business idea, how comfortable they are with technology, and what they want from the session. A second AI agent plays that person over a whole session, turn by turn, and the full conversation is saved.
One made-up entrepreneur
Who they are. Wants to run a mobile dog-grooming van in the east end of the city. Comfortable on a phone, less so on a laptop. Has a cousin who says the idea is a winner.
What they want from the session. To be told what to charge per groom.
A reply that fails. The coach names a figure. It may be a sensible figure. It still fails question two, because the entrepreneur did not reach it.
A reply that passes. The coach asks what nearby groomers charge, what a visit costs in fuel and time, and what the cousin's friends would pay. The figure comes out of that, and it is the entrepreneur's own.
Some of the made-up entrepreneurs are difficult on purpose. One wants the AI coach to write the business plan for them. One pushes for a yes on an idea that does not add up. One keeps changing the subject. None of them are real. No entrepreneur's data from any cohort is used in testing, which is the answer we would want from a vendor too.
Three readers then grade each saved conversation, starting with the cheapest.
- Plain code handles anything that needs no judgment: whether the coach saved the answer to the right milestone, whether it stayed in the right language, whether the session ended cleanly.
- A second AI reads the whole conversation and answers the five questions, one at a time, with a yes or no and a sentence of reasoning for each.
- A person reads a sample of conversations without knowing which version of the AI coach produced them, so they cannot favour the version they expect to be better.
Every conversation goes through the code and the second AI. The person's grades are there to keep the second AI honest, and any failure they find becomes a new made-up entrepreneur. The person is not a step we plan to remove once the second AI has earned our trust. A sample of every run is read by a person for as long as the AI coach is in front of cohorts.
Trusting a second AI to grade the first
A second AI reading the conversation is one more thing that can be wrong, and we did not want to find that out from a cohort. So before it graded anything that mattered, a person graded a set of conversations by hand, and we compared the two sets, question by question.
The usual advice is to adjust the grader until it agrees with the person. We went the other way more often than not: the disagreements usually turned out to be our fault rather than the AI's. When the person and the AI split on a reply, the question was vague enough that both answers were defensible, and rewriting the question settled it. A question is only allowed to decide whether a change goes out once the AI and the person agree on it across that set. The questions that have not reached that point stay with a person.
What changed for us
The first thing that changed is that an edit to the coach's instructions now produces a number. Each run reports how many conversations passed each of the five questions, next to the same figures from the run before. An edit that lifts question two and drops question four shows up as exactly that, and we can decide whether the trade is worth it before anyone in a cohort meets it.
The second is that trying a new AI model stopped being a debate. The same made-up entrepreneurs run against the current model and the new one, and the two sets of scores sit side by side. We still read the conversations. We no longer argue from memory about them.
The third is the one we care about most. When a coach in a program reports a session that went wrong, we write a made-up entrepreneur who reproduces it, and that entrepreneur joins the set. The same mistake cannot come back unnoticed, because every future change is tested against it.
The useful question is not whether the coach is good. It is whether this change made it better or worse, on questions two people would answer the same way.
We built this because we needed it. The AI coach is in front of real cohorts, and the programs running it did not sign up to be our testers. We work with real business owners and their coaches to get feedback on how the AI coach behaves. The checklist started from the tester's notes and now grows from that feedback, and every change has to earn its way past it before it reaches them.
What it does not tell us
We would rather name the limits than have you find them.
- It grades conversations, not outcomes. A high score means the business coaching behaved the way a good coach would. Whether the entrepreneur finished the program, or whether the business works a year on, are different questions, and this does not answer them.
- It cannot hear the session. A voice session has pauses, interruptions and a pace that a written transcript loses. How the voice feels is still tested by a person, with a headset on.
- Made-up entrepreneurs are tidier than real ones. They stay in character and answer the question. Real people bring a child into the room and come back to a different topic. The sample a person reads, and the feedback from real business owners and their coaches, are where we catch what the made-up ones miss.
Questions to ask anyone selling you an AI coach
If a vendor tells you their coach is tested, these four questions tell you what that means. None of them need a technical reviewer.
- Ask what is checked before a change reaches a cohort, and to see the score from the last run. A pass rate from a real run is a different answer from a description of a process.
- Ask who wrote the definition of good coaching. You want to see the list of questions a reply is graded on, and to know whether anyone who has coached a cohort helped write it.
- Ask what happens when one of your coaches reports a bad session. The answer you want is that it becomes a test the coach has to pass from then on.
- Ask whether your entrepreneurs' conversations are used in testing. If the answer is yes, ask how they are protected. If the answer is no, ask what is used instead.
If you run a small business program and want to see how the AI coach works with a cohort, and what one of these test runs looks like, you can book a demo.
Sources
- Anthropic, Model deprecations.
- Anthropic, Claude Fable 5.1 migration guide.
- Lingjiao Chen, Matei Zaharia and James Zou, How is ChatGPT's behavior changing over time?, 2023.
- Anthropic, Demystifying evals for AI agents.