Banking QA’s new reality: Is point-in-time testing dead?

London-based Simon Hull

Banks, insurers and asset managers are rapidly discovering that deploying AI inside regulated environments is exposing weaknesses not just in technology stacks, but in the way financial institutions approach software testing, governance and operational resilience altogether.

As AI systems become increasingly agentic, adaptive and embedded into core business processes, traditional models of QA and validation are coming under mounting pressure.

Periodic audits, static controls and release-stage testing are proving insufficient for systems capable of changing behaviour dynamically, interacting autonomously and operating across fragmented enterprise environments.

The challenge facing financial services firms is no longer simply whether AI works, but whether organisations can continuously observe, govern and trust what those systems are doing once they enter production.

That shift is forcing testing teams to rethink long-established approaches to validation, risk management and oversight.

Increasingly, quality engineering leaders are being asked to move beyond release assurance and into continuous operational monitoring, explainability and runtime governance.

Simon Hull

According to Simon Hull, head of financial services and insurance at CreateFuture and formerly with Barclays, BlackRock, UBS and DBS, many firms are now entering “harder territory” as AI adoption matures across banking, insurance and asset management.

Hull wrote in a recent LinkedIn article that “the industry is moving beyond early rollouts and into harder territory,” adding that “the question is no longer what AI can do, but what organisations need to become to do it well.”

Hull pointed to growing recognition across the industry that AI-driven financial systems require fundamentally different engineering disciplines.

Referencing comments from Marnie McCormack, Managing Director and Head of Engineering Practices at JPMorganChase, Hull wrote that “the discipline required is systems engineering, not software engineering. Code is not the asset. The system the business can trust is the asset.”

For QA and software testing teams, that distinction carries significant implications.

“In an agentic world, where AI increasingly takes actions rather than simply responds to queries, that distinction matters,” Hull wrote.

He noted that McCormack highlighted several emerging engineering pressures, including the fact that “observability is not yet mature enough,” while “team context must be made explicit for agents,” and “audit constraints, including SOX obligations, need to be accessible to AI systems.”

Most critically for financial testing environments, Hull stressed that “testing must shift from point-in-time validation to continuous monitoring.”

The statement reflects a growing concern across the sector that AI systems cannot be treated like conventional software releases validated at a single moment in time. Instead, firms increasingly require persistent oversight of behaviour, outputs, controls and model interactions after deployment.

Hull added that “architecture needs more rigour up front, rather than being treated as something to resolve later.”

Continuous oversight

The convergence between testing, governance and operational resilience emerged as another major theme.

Hull referenced remarks from Rebecca Mackenzie, Head of Technology Risk and Control at Monzo Bank, who argued that “retrospective audit and periodic review are not enough.”

According to Hull, “the direction of travel must be toward continuous monitoring, adaptive controls and real-time oversight.”

The implications are significant for banks attempting to satisfy both innovation ambitions and regulatory expectations around explainability, accountability and resilience.


“Ownership becomes fragmented, decisions become automated, and visibility is reduced.”

– Simon Hull

Hull added that “traditional control environments often rely on people’s judgment as an implicit backstop. AI removes some of that human backstop, so organisations need to make the implicit explicit by surfacing, documenting and designing the judgements that were previously assumed.”

He also highlighted increasing pressure around explainability and traceability, particularly as firms prepare for stricter AI governance frameworks including the EU AI Act.

Hull pointed to Mary Drabble, where “96% of staff [are] using Copilot daily and AI embedded in a new contact centre solution.”

Viable ideas to be tested

Beyond technical architecture, Hull argued that many financial institutions still struggle to create environments capable of sustaining meaningful experimentation.

Peter Crouch

He referenced comments from Peter Crouch, Group Innovation Director at Lloyds Banking Group, who said “large banks and insurers do not lack ideas.”

Instead, “what they often lack is an environment where early-stage ideas can survive long enough to be tested.”

The observation highlights the growing importance of experimentation infrastructure, safe testing environments and rapid feedback cycles capable of supporting AI iteration without collapsing under governance friction or legacy processes.

Hull wrote that Crouch described “learning speed as the new competitive advantage,” while arguing that “leadership behaviour” often becomes “the bottleneck.”

“Innovation requires decisions with incomplete information and organisations that follow evidence rather than politics, legacy assumptions or sunk-cost roadmaps,” Hull added.

Legacy systems

Hull also challenged the common assumption that legacy infrastructure is purely an obstacle to AI transformation.

“In my own talk I wanted to challenge the assumption that legacy is only a constraint,” he wrote.

“While legacy systems do represent a large impediment to change, they are also one of the most significant strategic assets available to incumbents.”

According to Hull, incumbent institutions possess “deep reserves of proprietary data, transactional history, customer behaviour, risk outcomes and market intelligence, accumulated over years.”

The challenge for banks, he argued, is finding ways to unlock that data safely without becoming trapped by the underlying complexity of legacy estates.

Hull stated that “the organisations best placed to win with AI may not have the newest infrastructure, but they can unlock proprietary data without being held hostage by the systems that contain it.”

He described this as “leapfrogging the legacy debt trap through AI-first engineering.”

Hull concluded that financial services firms risk focusing too heavily on AI deployment activity while underestimating the structural change required to support trustworthy AI operations at scale.

“Organisations need to think in systems, not just tools,” he stressed, adding that “activity is not the same as readiness,” warning that long-term value will likely favour firms capable of building continuous governance, adaptive monitoring and resilient testing disciplines around AI systems rather than simply accelerating deployment.


16 SEPTEMBER IN LONDON

REGISTER TODAY – SIMPLY CLICK HERE


Why not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

SIGN UP HERE TODAY


REGULATION & COMPLIANCE

Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.


READ MORE


WATCH NOW


QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS