The most uncomfortable meeting I have sat in this year lasted eleven minutes. A capable team walked through their AI assistant release: accuracy up, hallucination rate down, latency cut by a third, every regression suite green. Then someone from finance asked what I would call the only question that matters: “So what changed in the numbers?” Nobody had an answer. Not a bad answer. No answer. The metric had never been named, so nobody had been watching it.
Three weeks later they found it. Support ticket volume was down, which everyone had celebrated, but renewals in one customer segment had softened. The assistant was answering configuration questions confidently and incorrectly for a specific product tier. Customers did not open tickets. They followed the wrong instructions, got a broken setup, and quietly decided the product was harder than it looked. The feature had passed every technical gate and was actively eroding the commercial position.
This is not a rare story. It is the default outcome of a process that measures models the way engineers measure models and then hopes the money follows.
The retesting treadmill is not a strategy
Here is the loop most teams are stuck in. A new model version drops. Someone reruns the evaluation set. Scores hold or improve. The release ships. Six to ten weeks later, another model drops, and the loop restarts. The work is real and the discipline is genuine, but the loop has no exit and no commercial anchor. It measures whether the system is still working the way it worked before, not whether it is worth anything.
I have seen the same pattern in a different costume. When we built consumer hardware in China with the Jean-Michel Jarre venture, we went from one product to eight over a few years. Every unit passed factory QA. Drop tests, thermal cycles, acoustic benchmarks, the lot. QA told us the product met specification. It told us nothing about whether the specification was the right one for a buyer standing in a store comparing two boxes. Those are separate documents, written by separate people, and only one of them determines whether you get a second production run.
Software teams lost this instinct because with traditional software the two documents overlapped enough to be treated as one. If the feature works, it works. With AI systems the overlap collapses. A model can be measurably better on your evaluation set and commercially worse in production, because the failure modes that hurt you commercially are rarely the ones your evaluation set contains. Your evaluation set contains the questions you thought of. Your customers ask the questions you did not.
The treadmill also creates a quiet organisational problem. When the acceptance criteria are technical, the owner is technical. Nobody in the commercial organisation has signed anything, so when revenue moves, nobody in the commercial organisation feels accountable for explaining it. You end up with a feature that has an engineering owner and no business owner, which is the standard configuration for something that never gets a second budget cycle.
Two documents, and only one buys you next year’s budget
The technical acceptance document is familiar: accuracy thresholds, latency ceilings, cost per thousand calls, safety filters, regression coverage. It is necessary. It is also, on its own, a document about whether you built the thing correctly.
The commercial acceptance document answers a different question, and it fits on one page. Which revenue, retention or cost-to-serve number should this release move, by roughly how much, over what period, and whose name is against it? That last part is not bureaucratic. A metric without a named owner is a metric nobody checks on a Monday morning.
Running a climate-tech company has made me blunt about this. Our customers buy an outcome, and when we propose something new, the first question internally is which line it touches. Does it shorten the sales cycle, raise the contract value, reduce the hours we spend delivering, or increase the chance the customer renews? If the honest answer is “it makes us more modern”, we do not do it. Not because modernity is worthless, but because a project with no named line item has no defence when the budget tightens, and budgets always tighten.
When I ran digital transformation for a private healthcare chain expanding across mainland China, the projects that survived internal review were the ones tied to something a regional manager already reported on. Appointment no-show rate. Time from arrival to consultation. Repeat visit rate within twelve months. Those numbers already had owners, already had history, already appeared in meetings that had nothing to do with technology. Attaching a new system to an existing reported number is the cheapest way to make it accountable, because the reporting infrastructure and the political ownership already exist.
The practical test I use now: before an AI release, ask the team to write the sentence “if this works, [named metric] will move from X to Y by [date], and [named person] will report it.” If the team cannot complete that sentence without hedging into vague language, the release is not ready, no matter how good the evaluation scores are. The model may be fine. The business case is missing.
Draw the trust boundary before you ship, not after the incident
The second thing that never appears in technical acceptance criteria is the handback. Where, exactly, does the agent stop and a human take over?
Most teams answer this with confidence thresholds, which is a reasonable engineering answer and an incomplete commercial one. Confidence is a property of the model. The trust boundary should be a property of the consequence. I care far less about how sure the model is than about what happens to the customer if it is wrong.
Draw the boundary by consequence class. Anything that touches price, contractual terms, regulatory statements, safety instructions, or an irreversible action on the customer’s account sits outside the boundary. The agent can draft, retrieve, summarise and recommend inside that zone, but a human confirms before it reaches the customer. Anything reversible, low stakes and cheap to correct sits inside. This is not a sophisticated framework. It is the same logic banks apply to transaction limits, and I learned it working on machine learning lending and personalisation at HSBC in Hong Kong, where the model’s output was never the last step in anything that mattered. The model ranked, scored and surfaced. A defined process decided.
The mistake I see repeated is treating the trust boundary as a safety feature. It is a revenue feature. Customers forgive slow. They forgive “I need to check that for you.” What they do not forgive is confident and wrong, because confident and wrong makes them question everything else you have told them. One bad answer about a configuration detail does not just cost you that interaction. It costs you the customer’s willingness to trust the next twenty answers, which is the entire value proposition of the assistant in the first place.
So write the boundary into the release document in plain language, and write what the handback looks like from the customer’s side. Does it feel like help arriving, or like the system failing? That difference is design work, and it is worth more than another two points of accuracy.
Canary it like a pricing change, because that is what it is
No competent commercial team changes prices for every customer on the same morning. You test a segment, hold a control, watch the conversion and churn signals, and roll forward or back. Model releases deserve identical treatment, and almost never get it, because they arrive through the deployment pipeline rather than the commercial one.
Ship the new model to a slice of traffic. Hold a genuine control group on the previous version for long enough to see the commercial signal, not just the technical one. That usually means weeks rather than days, because churn and renewal effects lag. Accuracy shows up in an afternoon. Trust damage shows up at the next renewal conversation.
Watch the leading indicators that sit between the two. Escalation rate to human agents. Repeat contact within seventy-two hours on the same issue. Task abandonment mid-flow. Silent drop-off, which is the dangerous one, because a customer who abandons quietly generates no ticket and no complaint and looks, on a support dashboard, exactly like a customer you helped. The team I mentioned at the start had a falling ticket count and read it as success. Falling ticket volume with flat or worse retention is not efficiency. It is customers giving up.
Keep the rollback path warm. The reason model rollbacks feel dramatic is that teams treat them as admissions of failure rather than as the normal operation of a controlled release. A pricing experiment that gets reversed is a good experiment. So is a model version that gets pulled at eight percent of traffic instead of a hundred.
What to change before the next release
Three things, and none of them require new tooling.
Write the commercial acceptance document alongside the technical one, one page, with a named metric and a named owner from outside engineering. If the metric already appears in an existing management report, better still. Define the trust boundary by consequence rather than by model confidence, and describe the handback as an experience, not an error state. Then release to a slice with a real control group and hold it long enough for retention signals to surface.
The question “is the model good enough?” has no terminal answer, which is exactly why it keeps you on the treadmill. The next model will always be better on something. Replace it with a question that terminates: which commercial number does this release move, and who reports it? Teams that can answer that get a second budget cycle. Teams that cannot get a very short meeting.
I write from twenty years of building businesses between Europe and Asia. If your company is facing this, start a conversation.