
This is a #paidpartnership with Commonwealth Bank.
The hardest part of enterprise AI was never making the model smarter.
Most teams are still trying to win a race that’s already been decided. They wait for the next model, whether that’s a higher benchmark, a longer context window, a cheaper token, as if the distance between a promising demo and a system customers depend on were a capability gap. It isn’t. It’s a trust gap. And there’s a number that proves it.
On our latest episode of What The Tech, we sat down with Beibei Guo, Distinguished Engineer at CommBank, the bank that just ranked #4 in the world and #1 in Asia-Pacific in the 2025 Evident AI Index. Twenty-plus years across systems engineering and data platforms. She’s watched the internet, virtualisation, and cloud all arrive with the same promise and the same trap. Her read on AI is unsentimental: the model getting smarter was never the bottleneck.
We’re not the only ones finding this. MIT’s NANDA initiative, in The GenAI Divide: State of AI in Business 2025, reported that 95% of enterprise generative-AI pilots fail to move the P&L and the cause isn’t model quality. It’s what MIT calls the “learning gap”: the failure to wire AI into the workflows, structures, and habits of the organisation. The number that should frighten you sits next to it: the share of companies abandoning most of their AI initiatives jumped from 17% to 42% in a single year.
The conventional wisdom
The default mental model says AI value is a function of model intelligence. Pick the smartest model, point it at the problem, wait for it to be “good enough.” Under that logic, every failed pilot is a model problem and every model problem gets solved by the next release.
It’s a comfortable story, because it asks nothing of the organisation. You just wait.
Why it’s incomplete
Beibei has a sharper frame. AI without trust, she says, is like a talking dog: a delightful novelty you don’t quite know what to do with. Impressive, briefly. Useless in production. What turns the talking dog into something you’d deploy, her words is trust: “with trust, it’s like Mr. Data in Star Trek.” Still a little strange, but dependable enough to bring along on the journey.
That is exactly what MIT found, stated in the language of P&L. The 95% don’t fail because the model isn’t clever enough. They fail because the cleverness never became something anyone could rely on. Intelligence was abundant. Trust was missing.
What trust actually looks like
Trust isn’t a feeling you wait to develop. It’s a property you engineer into the workflow. Three moves from the conversation:
Put the gate in the workflow, not the prompt. CommBank is an early adopter of the AWS DevOps Agent, announced at re:Invent 2025. When an incident fires, it correlates logs, telemetry, code, and deployment data and hands the on-call engineer a root-cause hypothesis, work that used to mean 40 minutes of dashboard-spelunking before anyone could even think. What it does not do is fix anything. As Beibei put it, she’s glad it does what she told it to do and not go ahead and fix it. The human approves the action. “If it gives you a step that drops your production database,” she said, “it relies on you to tell it not to.” Accountability never leaves the person.
Shape what you can’t define. You can’t specify exactly what a model will output. You can shape it and you have to understand the shape well enough to know where it’s safe to point. That understanding is the work. It doesn’t arrive in a release note.
Keep a human on the decision that carries weight. Beibei described a small production change like adjusting a system parameter with real downstream consequences. She used Claude to draft the test cases and run them; she made the call and stayed accountable for it. The AI took the repeatable work. The judgment stayed with her, on purpose. “I’d like to be accountable in that way,” she said. “Selectively.”
What doesn’t get easier
Here’s the hard edge. When a model can do more, the few things it can’t do matter more, not less. Accountability doesn’t compress. Judgment doesn’t compress.
The MIT data has a tell buried in it: pilots that paired internal specialists with outside expertise hit a 67% success rate; IT-only builds managed 22%. The differentiator wasn’t a smarter model. It was people who owned the outcome.
Beibei’s version is human-scale. Her graduate engineers shipped production-grade data-comparison systems across hundreds of terabytes within months, not because she handed them a prompt, but because she handed them “a few grads, a few models, and a principal engineer” who met them weekly and owned the result.
The delivery
So here’s the whole thing, paid off: the smartest model on the leaderboard will not move you from the 95% to the 5%. Trust will. The teams crossing the divide aren’t the ones who waited for a better model, they’re the ones who built the gate into the workflow, kept a human on the consequential call, and treated accountability as the part that never gets automated.
Your AI doesn’t have a smartness problem. It has a trust problem and unlike the model, that one is yours to fix.
Listen to the full episode
In our conversation with Beibei Guo, we get into:
Why “demo is easy, and nothing is easy inside the corporate world”
The talking-dog-vs-Mr-Data test for any AI system
How CommBank lets agents do the heavy lifting while accountability stays human
Why the engineers who win pair models with mentorship, not prompts
The trade-off every leader is now navigating: the risk of adopting vs. the risk of not adopting
🎧 Listen on Spotify · Apple Podcasts · Amazon Music · YouTube
📬 Subscribe on Substack for episode breakdowns and behind-the-scenes thinking.
Question for readers
If you pulled the smartest model out of your stack tomorrow and dropped in the second-smartest, what would actually break? If the honest answer is “not much,” your bottleneck was never intelligence.
Sources:
MIT NANDA, “The GenAI Divide: State of AI in Business 2025”;
AWS, “From AI agent prototype to product: Lessons from building AWS DevOps Agent” (re:Invent 2025);
Evident AI Index 2025.
Episode quotes from What The Tech (AU), our conversation with Beibei Guo.

