This is the second article in QA Financial’s two-part series examining how Barclays is preparing autonomous AI agents for deployment inside one of the world’s most highly regulated banking environments.
Part one explored testing, telemetry and production controls, while this feature examines why software assurance is becoming one of the biggest challenges facing enterprise AI.
For the past three years, enterprise AI has largely been measured by how quickly organisations could build something.
Banks are beginning to ask a different question. Can they prove it is safe enough to deploy? That shift may prove to be one of the biggest changes facing quality engineering since DevOps first brought testing into continuous delivery.
As autonomous AI systems move from internal pilots towards production workloads, software testing is expanding beyond validating functionality into something much broader: producing evidence that complex AI systems can be trusted.
Andy McMahon, Principal AI Engineer at Barclays, believes too many organisations remain focused on getting agents into production before building the engineering disciplines needed to support them.
Speaking on The Brave Technologist podcast last month, he suggested the industry’s biggest challenge is no longer capability but production readiness.
One of the strongest moments in McMahon’s interview came when he described conversations after one of his conference presentations.
“It was interesting a lot of people came up to me after my talk and says, ‘Oh, we’re now we’ve deployed a lot of agents… we’re now starting to think about observability.'”

For software testers, that order of events should feel uncomfortable. Observability has traditionally been designed into distributed systems from the outset, providing engineers with logs, traces and metrics before production traffic ever arrives.
Yet McMahon suggested some organisations are reversing that sequence with AI. “They’ve been in the race to innovate like you said into like I can do this thing with agents and it’s going to drive like this benefit…”
“…but I’m not actually tracking what it’s doing … or actually my eval quite thin … I have kind of not done a really rigorous batch of testing on it because the race was to get something out there,” he stated.
It is perhaps the clearest warning in the entire interview. The challenge is not that organisations lack AI capability. It is that many are building production systems before they have built production assurance.
That observation echoes a wider trend across financial services. The Financial Conduct Authority has increasingly argued that AI assurance extends well beyond model validation, encompassing governance, deployment context, human oversight, evaluation techniques and operational controls.
In other words, software testing is moving upstream and downstream at the same time.
From shipping software to proving value
McMahon also believes the industry is asking the wrong questions. “I think a developer and engineer scientist sort of mind will always ask can I build this? But they should ask before that should I build this?”
It is a deceptively simple observation. Generative AI has dramatically lowered the cost of building prototypes. Creating another chatbot, another assistant or another Jira automation is easier than ever.
The harder question is whether any of those systems solve meaningful business problems. Everyone’s like, ‘Oh, I can build that. I can build a bot that reads Jira tickets’,” he said.
“But what you end up with is like a 100 bots that read Jira tickets. Is that really the best use of your time?”

For McMahon, engineering success is becoming less about technical possibility and more about measurable value. “Where’s the value?”
He argued organisations often discover the most valuable opportunities are also the hardest engineering problems. For QA teams, that changes the role of testing.
Testing no longer supports software delivery alone. Increasingly, it provides confidence that investment decisions are producing resilient, governable systems capable of delivering sustained business value.
Why banks start internally
Barclays’ own deployment strategy reflects that philosophy. Rather than immediately exposing customers to autonomous AI, McMahon said internal use cases provide the safest place to learn.
“Internal use is always going to be where you start. It’s the easiest place to play. It’s the least risky,” he pointed out. “We are now starting to push some stuff out customer-facing which is really exciting.”
That approach mirrors the crawl-walk-run strategy now adopted by many financial institutions. “We’re very much taking an approach… where it is always crawl walk run,” he continued. “It’s very much things in our environment are quite bounded.”

Rather than pursuing unrestricted autonomy, the bank is gradually expanding capabilities while retaining tightly controlled permission structures and governance.
For software testing professionals, the lesson is clear. Production confidence is accumulated rather than declared. Each successful deployment becomes evidence for the next.
McMahon also cautioned against treating today’s AI boom as unprecedented. “I’ve seen all this before.”
He compared today’s excitement around agentic AI with earlier waves of machine learning and cloud computing. “There was a huge race to innovate… And then what happens is big rush and then we sort of acclimatise and equilibrate.”
Only then do organisations begin asking the questions quality engineers have often been asking from the beginning. “Are you doing things actually scalably? Are you doing things actually kind of in this controlled way? Are you monitoring the cost?” he mentioned.
Those questions sound remarkably familiar because they have long been central to mature software engineering. The difference today is that autonomous systems make the answers even more consequential.
QA’s expanding role
Taken together, McMahon’s observations point towards a future in which quality engineering becomes even more central to enterprise AI. Banks are unlikely to compete solely on model performance.
Instead, competitive advantage may increasingly depend on how confidently institutions can deploy AI into production while satisfying regulators, auditors and customers that those systems remain observable, governed and resilient over time.
That aligns closely with Barclays’ own engineering journey. The bank has already invested heavily in shift-left testing, observability and automation. McMahon’s vision suggests those capabilities now form the foundation for the next generation of AI assurance.
Perhaps the most important takeaway from the interview is that autonomous AI does not reduce the need for software testing.
It dramatically expands it. Telemetry becomes evidence, while observability becomes assurance and evaluation becomes continuous rather than occasional. And production is no longer the end of testing. It becomes the place where testing truly begins.
This was the second article in QA Financial’s two-part series examining how Barclays is preparing autonomous AI agents for deployment inside one of the world’s most highly regulated banking environments. Part one explored testing, telemetry and production controls.
NEXT MONTH



REGISTER TODAY – SIMPLY CLICK HERE
Why not become a QA Financial subscriber?
It’s entirely FREE
* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *
REGULATION & COMPLIANCE
Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.
READ MORE
- Goldman puts AI coding to the test
- How to test AI models that banks do not control
- OpenAI, Filigran and SunTec: the latest vendor and product news
- Sygnum: Testing AI is ‘a measurement problem’
- Banks’ ‘code for all’ push raises testing risks
QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS



