For program leaders · Proactivly AI · August 7, 2026 · 7 min read
What the OECD found about AI tutoring — and what it means for your cohort
An AI built to tutor raised practice performance by 127%. A standard answer-giving chatbot raised it by 48% — and once it was taken away, its users scored 17% below students who had never used AI at all. The gap is a design decision, and it is one you can check for before a tool reaches your cohort.
Most of the AI conversations we have with program directors get stuck in the same place: allow it, or ban it. Your coaches suspect entrepreneurs are already using it. Your funders want to know your position. And the honest answer — that it depends on the tool — sounds like ducking the question.
There is now a piece of research that makes it less like ducking. It suggests allow-or-ban is the wrong axis entirely, and that the thing worth having a policy about is which kind of AI is in front of your cohort.
What the study measured
OECD Digital Education Outlook 2026: in a field experiment with high-school maths students, an AI built to tutor raised practice performance by 127%, against 48% for a standard answer-giving chatbot. But once access was taken away, the students who'd used the standard chatbot scored 17% worse on a closed-book exam than students who had never used AI at all.
Three figures, and they measure three different things:
- 48% — how much better students did during practice with a standard chatbot interface.
- 127% — how much better they did during practice with a version engineered to support the learning process rather than to hand over answers.
- −17% — how the standard-chatbot group did on a closed-book exam after access was removed, compared with students who never had AI at all.
Note what that last figure covers and what it does not. It is the standard-chatbot group against the never-exposed group. Don't read it as a verdict on the tutoring version — that is not what the number says, and we are not going to fill in a comparison the source doesn't state.
The number that should shape your policy is the negative one
The two practice figures are the ones that get quoted, and they are the less interesting pair. Of course people do better on a task while holding a tool that does the task.
The finding with consequences is that a group ended up worse than if they had never used AI at all — not while using it, but afterwards, on their own.
For an entrepreneurship program that should land uncomfortably close to home, because your program is the practice phase. Twelve weeks with coaches, structure, deadlines, and something helpful open in a browser tab. The exam is everything after. Nobody in your cohort is graded on how well they perform while the tool is open; they are graded, eventually, by whether the business works when the program has ended and nobody is checking in.
So the question a director actually needs answered is not "did the cohort produce better plans this term." It is: can each entrepreneur explain and run the plan when the scaffolding comes off. Those two things can move in opposite directions, and the OECD's data is a case where they did.
The variable was design, not AI
The two groups in that experiment were not AI versus no-AI. They were the same underlying technology, presented two ways: a standard interface that answers, and a version engineered to support learning.
That is the part worth carrying into your next vendor conversation. "Are you using AI" tells you almost nothing about whether a tool helps a cohort. It has roughly the information content of asking whether a program uses spreadsheets. What decides the outcome is what the tool does when someone brings it a hard question — and that is a design decision the vendor made, on purpose, and can tell you about.
Four things to check before a tool reaches your cohort
None of these need a technical evaluation. A coach can check all four in twenty minutes.
- Does it ask before it answers? Bring it a vague idea — "I want to start a bakery" — and see what comes back. A business plan means it is an answer-giving tool. Two or three questions about who the customer is means it is doing the thing the tutoring version did.
- Can a coach see the thinking, or only the output? A polished document tells a coach nothing about where the entrepreneur actually is. If the tool surfaces the reasoning behind a milestone, the coaching conversation starts somewhere useful instead of starting from a document the coach has to reverse-engineer.
- Does what comes out sound like the entrepreneur? Read three outputs from three different people in the cohort. If you cannot tell them apart, the tool wrote them, and the entrepreneur will not be able to defend a word of it in front of a funder.
- What survives removal? Ask an entrepreneur to talk through their plan with the laptop shut. This is the closed-book exam, and it is the only one of the four checks that measures what the OECD result is actually about.
What this study does not tell you
We would rather state the limits than have you find them:
- It studied high-school mathematics students, not entrepreneurs. Different domain, different age, different stakes, and maths has right answers in a way that a go-to-market decision does not.
- It is one field experiment. It is well-documented and published by a serious institution, and it is still one result.
- The practice gains measure performance while using the tool, not learning that stuck. Anyone quoting "127% better" as though it means students learned more than twice as much — including anyone quoting it at you — has skipped a step.
What the study supports is narrow and still useful: how an AI is designed changed the outcome more than whether AI was present, and the harm showed up only after access ended. That is enough to change how you evaluate a tool. It is not enough to predict what will happen in your cohort.
Where we stand
This finding is the argument Proactivly is built on, so treat the rest of this section as interested rather than neutral. Our position is that an AI which hands an entrepreneur a finished business plan has produced a document and taught nobody anything — and that the useful machine is the one that asks the questions a good coach would ask when there are more entrepreneurs than coaching hours.
Our own evidence is small and we will not oversell it: one partner program went from roughly 60% to 88% cohort completion on its first Proactivly cohort. That is one program and one cohort. Weigh it accordingly.
If you want to see what "asks before it answers" looks like in a real cohort, book a demo — we will take one milestone end to end and you can judge the design question for yourself.
Sharing this with your cohort? The entrepreneur-facing version of this argument is in our guide to AI-assisted business planning, written for the person doing the work rather than the person running the program.