Your AI Pilot Worked. That's Why It Stalled.
$30 million in open quotes, an AI pilot that looked like it was working, but no movement on the P&L. The model produced the work. The business never changed around it.
I sat down with $250 million logistics client looking at $30 million in open quotes.
They had already approved an AI pilot, and on the surface it looked like it was working. The model summarised tenders, drafted responses, pulled numbers out of spreadsheets, and surfaced patterns faster than the team could do by hand. If you watched the demo in isolation, you would have called it a success. But nothing moved on the P&L.
The easy answer was to blame the model and move on.

The model was doing what it had been asked to do, but the business around it had not changed enough to absorb the output. Quotes still waited for review, managers still checked the same steps, legal still touched work that should never have reached them, and sales still carried the same bottlenecks, just with better drafts sitting upstream.
That is the trap more executive teams are running into now. The first phase of AI was about proving the tools could produce something useful. The next phase is less glamorous and much more commercial. It is about deciding who checks the output, what gets sampled, what goes straight through, which task disappears, which cost line moves, where saved time goes, and who owns the error when something plausible but wrong gets through. If those answers are vague, the pilot does not scale. It just creates more work for the same people to review.
A recent paper from Hyunjin Kim, Dahyeon Kim and Rembrand Koning helps explain why. In Mapping AI into Production: A Field Experiment on Firm Performance, the researchers studied 515 high-growth startups and found that firms shown how other companies reorganised production around AI discovered 44% more AI use cases, completed 12% more tasks, were 18% more likely to acquire paying customers, and generated 1.9x higher revenue. Same technology, different result.
AI pilots are still built to answer a question that mattered two years ago: can the model produce the output? That was useful when the technology itself was uncertain. It is not enough now. In most white-collar workflows, AI can already produce something useful, whether that is a claims summary, a sales email, a contract comparison, a board pack draft, a policy check, or a research brief.
But my question is whether the organisation can stand behind the output once it exists. Can it be checked quickly enough to matter? Can it be trusted often enough to remove manual work? Can it be governed without turning managers into permanent quality control? Can it move a ledger line the CFO will recognise? Can it remove work, or does it just create more work for people upstream to inspect?
This is where the gap starts to open - the CEO sees speed, the CFO asks where the saving lands. The COO asks which workflow changes. The GC asks who owns the risk. The frontline asks whether they still have to check everything. Nobody has a clean answer, so the pilot stays alive in the worst possible state: not dead, not scaled, just drifting.....
I'm seeing this pattern more often now. The AI works well enough to create momentum, but the organisation has not redesigned the work around it. AI business cases make the same mistake here. A team says it can generate 1,000 customer summaries a day, which sounds impressive until you ask the right questions. Who checks them? What gets sampled? What gets escalated? What goes straight through? What error rate is acceptable? What happens when a plausible error gets through? Which human task disappears? Which cost line reduces? Where does the freed capacity go? If those questions are unanswered, the ROI is not real yet.
This is also how AI creates false confidence inside the business. The dashboard shows more output, the team feels faster, and the vendor declares the pilot a success, but managers are now reviewing twice as much work, legal gets dragged into edge cases, operations handles exceptions manually, and finance still cannot see the reduction because the old workflow stayed intact and AI just sat on top of it.

This is why I still come back to the same four commercial doors: revenue, cost, time, or risk. If a pilot does not increase revenue, reduce cost, compress time, or materially reduce risk, it should not survive the funding conversation. If it does pass that filter, the next question is simple: which line moves?
Not productivity....
Not efficiency......
Not better experience.....
Name the line on your P&L. A customer service summary tool should reduce handling time, improve first-contact resolution, cut after-call work, or lower complaint risk. A sales assistant should increase proposal volume, improve response speed, or free rep time for live opportunities. A contract review tool should reduce legal backlog, surface risk earlier, or compress approval cycles. If the use case cannot attach cleanly to one of those outcomes, it is not ready for production funding. It may still be interesting, but that is not enough.
Teams automate before they eliminate, and then never reallocate. AI does not create operating value until someone removes the old work. If the same report still gets written, reviewed, reformatted, checked, discussed, and approved by the same people, the business has not changed. It has simply added AI to the congestion. The order matters. First remove the step AI makes unnecessary. Then automate what still needs to happen. Then reallocate saved time and judgement to where it still earns its keep, whether that is approvals, exceptions, negotiations, customer moments, risk calls, commercial decisions, or coaching.
If AI saves a manager eight hours a week and those eight hours vanish into more meetings, the business got nothing. If those hours move into coaching the frontline, calling at-risk customers, or fixing a root-cause process issue, this is case that survives scrutiny.
A successful pilot often creates the next bottleneck, which is why human oversight is not a strategy in itself. It is a design choice. What gets checked? What gets sampled? What gets approved automatically? What gets escalated? Who owns the final call? What evidence is retained? Which decision rights change?
If you want a practical test before the next steering committee meeting, run one pilot through five questions.
- Which revenue, cost, time, or risk line is it meant to move?
- What baseline are you using?
- Which step disappears if the pilot works?
- What gets approved automatically, and what still needs review?
- Where does the saved capacity go if the pilot succeeds?
If the team cannot answer those in plain English, the pilot is not ready for scale.
That is also the point where I need to touch on governance - if the workflow touches customer outcomes, employee decisions, privacy-sensitive information, or a regulated approval path, the approval model has to be mapped before scale. Under the Privacy Act and current OAIC guidance on commercially available AI, the business deploying the workflow still owns the decision path, the recordkeeping, and the accountability if something goes wrong. You do not wait until rollout to ask who signs off, what gets retained, which exceptions become legally material, and what evidence the board will want when something inevitably goes wrong. See last weeks newsletter for more on this..

Before Friday, take one stalled AI pilot and write this sentence for your CFO:
"This pilot deserves production funding because it will improve [revenue/cost/time/risk line] from [current baseline] to [target outcome] by [date], with validation handled through [approval model], and saved capacity reallocated to [higher-value outcome]."
If you cannot write that sentence in under three minutes, do not scale it. Reset it, kill it, or redesign the workflow until the sentence becomes true.
AI can now produce the work, but will your org absorb it?
Ramon.
About Applied AI Australia
AI commentary often stops at the model. Applied AI Australia is for executives who need to turn AI into something a CFO, COO, or board can see in the P&L.
We help Australian companies between $100 million and $1 billion turn AI into revenue, margin, time back, and better operating discipline. One podcast and one newsletter each week, built so you can brief a board in under an hour.
Listen now and subscribe: Available on Apple and Spotify
If someone forwarded this email, subscribe to the newsletter: Here.
Disclaimer
Nothing in this newsletter is advice. It is research, pattern recognition, and practical operating observations organised for executives. Before acting on any of it, speak with your own adviser, or get in touch if you want to discuss an engagement.
Ready to deploy AI with confidence?
Get board-ready frameworks and strategic guidance for Australian executives navigating AI transformation.
Discuss your AI problem