The demo was excellent. The agent read the inbound order, checked stock across two warehouses, flagged a shortage, proposed a split shipment and drafted the customer email. Twelve seconds, end to end. Everyone in the room nodded.

Then the operations lead asked one question, and it had nothing to do with the model: “If this thing splits a shipment wrong on a Friday afternoon and the customer’s line stops on Monday, whose name is on that?”

Silence. Then the founder started talking about accuracy rates.

That was the moment the deal died, and I have now watched some version of it play out enough times, from both sides of the table, that I no longer treat it as bad luck. I sit on the buying side today as CEO of a climate-tech company, and I have sat on the building side for twenty years. The pattern is the same in manufacturing, in SaaS, in infrastructure. Technically superb AI products stall at the first serious commercial conversation, and the founders diagnose it as a capability gap when it is a liability gap.

The pattern, and the wrong fix

The response is almost always the same three moves. Add more AI language to the homepage. Optimise inference cost so the gross margin story looks cleaner. Start hunting for a moat, usually framed as proprietary data or fine-tuning.

None of those address what the buyer is actually doing in the room, which is risk arithmetic, not capability assessment.

Run the numbers the way an operations director runs them. Say the agent touches 4,000 order lines a month. At 99% accuracy, that is 40 wrong lines a month. The clerk it replaces is probably running at 97 or 98%, so on pure accuracy the machine wins. But the clerk’s errors are legible. They happen one at a time, they get caught by the next person in the chain, and there is a human who noticed, apologised and fixed it. The agent’s 40 errors can be correlated, because a single bad rule or a misread field produces the same mistake forty times before anyone sees it. And when the customer calls, there is no one to put on the phone.

The buyer is not comparing your model to a human. The buyer is comparing a distributed, absorbable, well-understood error profile to a concentrated, novel, unexplained one. The upside of the deal is maybe three to five percent of a cost line. The downside is a stockout, a missed shipment, an inventory write-off, or in regulated sectors an audit finding. That asymmetry is why “our model is better” lands with a thud.

I learned this version of the arithmetic at HSBC in Hong Kong, building machine learning into lending decisions. We had models that outperformed the incumbent scorecards. That was the easy part and it took a fraction of the time. The hard part, the part that consumed quarters, was model risk governance: what happens on a wrong decline, who reviews the override, how do you reconstruct the decision eighteen months later when someone asks. The institution was never going to let performance buy its way past accountability. Enterprise buyers of AI agents behave exactly the same way, just with less formal vocabulary for it.

Accountability, scope, reversibility

What unlocks the commercial conversation is not capability. It is three properties that have almost nothing to do with the model.

The first is accountability. Somebody has to own the outcome, and if the answer is “the AI” then no one does, which is unacceptable. The workable answer is that a named human owns the outcome and the product’s job is to make that ownership cheap. That means the audit trail is not a log file buried in an admin panel, it is a product surface. When the agent changed a record in the ERP, what did it see, what rule fired, what did it consider and reject, and which human had the chance to intervene. If an operations manager can reconstruct any single action in under a minute, she can defend it in a meeting, and if she can defend it she can approve the rollout.

The second is scope. The winning positioning is embarrassingly narrow. Not “AI for supply chain” but “we handle backorder allocation for distributors on this ERP, and here is the audit trail.” Narrow scope is not a lack of ambition, it is a way of bounding the blast radius. A buyer can reason about the worst case of one workflow. Nobody can reason about the worst case of a general agent with write access.

The third is reversibility, and this is where software founders undervalue what they already have. I spent years building consumer hardware in China with the Jean-Michel Jarre venture, growing one product into a line of eight. In hardware, a mistake ships in a container and there is no undo. You run pilot builds, you pay for tooling changes, you live with what you sent. Software agents can be built with rollback: every write reversible, every batch action undoable as a unit, a kill switch that any supervisor can hit without calling support. That is an enormous commercial asset and most teams never put it on a slide. Tell a buyer that any action taken in the last 72 hours can be reversed in one click, with the downstream effects listed, and you have removed more objection than any accuracy benchmark ever will.

Sell autonomy in stages, and price it that way

The other structural mistake is selling the destination instead of the path. Founders pitch the autonomous end state because that is where the value is, and buyers hear the maximum-risk configuration on day one.

Sell the ladder instead. Start in shadow mode, where the agent observes and proposes but changes nothing, running alongside the existing team for six to eight weeks. The deliverable at the end of that period is not a demo, it is a number: agreement rate between the agent and the humans, broken down by case type, with the disagreements analysed. That number does something no benchmark does, because it is measured on their data, in their mess, with their edge cases.

From there, move to suggest-and-approve, where a human clicks through. Then to execution inside thresholds, where the agent acts alone on the low-risk band, say orders below a value or within a defined SKU class, and escalates everything else. Only then, and only per case type, full autonomy inside a bounded scope with exception routing.

Price the stages differently. Shadow mode should be paid, modestly, because free pilots have no organisational weight and die quietly. Execution should cost more than suggestion, because it is worth more. This structure converts a binary trust decision into a sequence of small ones, and it gives the internal champion something to show his boss every few weeks. I ran digital transformation for a private healthcare chain expanding across mainland China, and the projects that succeeded were never the ones with the best feature set. They were the ones where the clinical and admin staff could see the thing working before they were asked to depend on it.

The moat is the integration depth, not the model

Founders worry about the wrong moat. The model layer is a commodity on a six-to-twelve month refresh cycle, and anything you fine-tune today is reproducible by someone else soon enough.

What is not reproducible is everything you accumulate by being inside one customer’s operations for eighteen months. The 200 exceptions in their warehouse management system that exist because of a decision someone made in 2014. The vocabulary their planners use that matches nothing in the schema. The escalation paths, the seasonal patterns, the three suppliers who always ship short. The trust record: eleven months of clean actions, two incidents handled well, a rollback that worked when it mattered.

A competitor with a better model still has to re-earn all of that, and the buyer knows the cost of re-earning it. When I co-founded one of China’s early location-based platforms and we passed 300,000 users, the asset was never the technology, which others copied inside a year. It was the accumulated behaviour and the position we held in people’s routines. Enterprise AI is the same shape. Integration depth and operational trust compound. Model quality does not.

What to change on Monday

Cut the homepage down to one workflow and say it in the language the buyer’s operations team uses, not the language of your engineering team. Put the audit trail in the demo, early, before the capability showcase. Write the blast radius into the contract, meaning the explicit list of what the agent can and cannot touch, and the reversal guarantee. Build the shadow-mode phase as a priced, productised first step with agreement rate as the deliverable. Stop describing the model.

The buyer is not asking whether your AI works. The buyer is asking what happens on the day it does not, and who will be standing there when it does. Answer that question well and the capability conversation takes care of itself.


I write from twenty years of building businesses between Europe and Asia. If your company is facing this, start a conversation.