UK plan broadens AI testing burden for banks

Starling Bank CIO Harriet Rees

Financial institutions may need to demonstrate not only that their AI models perform as expected, but also that the regulatory rules and internal policies governing their decisions have been translated into consistent, testable logic.

That is a central quality assurance implication of HM Treasury’s new Financial Services AI Adoption Plan, developed by Harriet Rees, group chief information officer at Starling Bank, and Rohit Dhawan, head of AI and advanced analytics at Lloyds Banking Group.

The plan calls for the creation of a voluntary, industry-led assurance scheme for third-party AI systems. It also recommends shared reporting of AI incidents and near misses, greater scrutiny of critical AI and cloud suppliers and cross-sector scenario testing.

Rohit Dhawan

Its publication comes as banks move beyond limited AI experiments and deploy the technology across a growing number of operational and customer-facing processes.

The Bank of England and Financial Conduct Authority found in a 2024 survey that 75% of responding financial services firms were already using AI. Firms expected their median number of AI use cases to rise from nine to 21 within three years.

The Treasury plan acknowledges that financial firms already operate within an extensive regulatory framework. The difficulty is translating that framework into practical controls that can be applied consistently to individual AI use cases.

“The core challenge now is not the absence of regulatory support, but its accessibility, consistency and practical application across the full breadth of the sector,” the plan states.

This creates two connected assurance problems. The first is determining whether an AI system operates reliably, securely and within agreed performance limits.

The second is testing whether the decisions it supports follow the institution’s approved interpretation of applicable regulations and policies.

Questions raised

In response to the plans, Calvern James, UK country manager at Rulemapping Group, argued on LinkedIn that the resulting governance challenge extends well beyond approving and monitoring the underlying models.

“AI governance cannot stop at the model,” James explained. “Financial institutions are investing heavily in model approval, testing, bias controls, monitoring and performance.”

Calvern James

He added: “But is the regulatory reasoning guiding the decisions those models influence being implemented with the same rigour?”

James stressed that “a firm can have strong controls around its models and still face a more fundamental question: How do regulatory principles and internal policies become clear, operational decision logic for the processes those models support?”

This distinction becomes increasingly important as AI is introduced into areas such as customer support, fraud detection, credit decisions, financial guidance and payments.

A technically accurate output could still produce a poor or non-compliant outcome if the governing policy has been interpreted incorrectly or applied inconsistently.

James identified three layers that firms may need to connect: “What the model can do”; “What the organisation permits it to do”; and “how the organisation demonstrates that decisions followed approved reasoning.”

“The first is primarily a technology challenge,” he continued “The second and third are where regulatory implementation becomes operational governance.”

Testing may therefore need to cover the decision logic surrounding a model, as well as the model itself. That includes establishing which rule was applied, how it was interpreted, what evidence informed the decision, where human judgement entered the process and who remained accountable for the outcome.

It also raises a consistency question: whether equivalent cases processed through different teams, systems or distribution channels would receive the same treatment under the firm’s approved policies.

“And they do not get easier as AI scales,” James warned. “More use cases mean more decisions, made faster and across more of the business.”

Third-party AI

The Treasury plan also proposes an AI Third Party Assurance Framework through which qualified assessors could evaluate external models and providers against agreed standards covering areas such as reliability, cybersecurity, transparency and data protection.

The framework could eventually operate as a standardised audit protocol similar to SOC 2. Regulators might accept certification as evidence that a bank has completed part of its baseline due diligence, reducing repeated assessments of the same provider by different firms.

Certification would not, however, remove the responsibility of individual institutions to manage the risks created when a model is implemented within their own environment.

This distinction is important for QA teams. A third-party model may satisfy a common assurance standard but behave differently when connected to a particular bank’s data, systems, workflows and customer journeys. Firms would still need integration, security, performance, resilience and policy-conformance testing appropriate to each deployment.

The proposed framework could nevertheless establish a more consistent starting point for AI procurement and testing, particularly where multiple banks currently conduct similar assessments of the same model or technology provider.

Testing incidents and near misses

Another recommendation is the establishment of a voluntary, industry-wide repository for AI incidents and near misses.

Such a repository could help banks convert production failures and unexpected model behaviour into reusable test scenarios. An incident experienced by one institution could inform regression testing, monitoring rules and risk assessments elsewhere in the sector without identifying the firm involved.

The plan also recommends assessing important AI and cloud suppliers under the UK Critical Third Parties regime. This would place greater emphasis on testing concentration risk, service disruption, data security and the ability of banks to continue operating when a major external provider fails.

HM Treasury additionally calls for cross-sector scenario testing covering AI-related disruption across interconnected industries including financial services, energy, telecommunications and cloud infrastructure.

These exercises would test whether existing contingency plans and incident-response arrangements remain adequate as financial institutions become more dependent on shared AI infrastructure.

Agentic payments raise the stakes

The plan’s focus on agentic payments introduces a further testing challenge. Autonomous agents capable of initiating or managing transactions will require controls extending beyond conventional model accuracy.

Banks may need to test how an agent’s identity is established, whether customer consent remains valid, which transactions it is authorised to execute and what happens when several agents or third parties interact across the same payment chain.

The Treasury recommends a trust framework built around legal accountability, Know Your Agent protocols and machine-to-machine authentication.

For testing teams, this is likely to require repeatable validation of transaction limits, authentication, fraud controls, exception handling, revocation and human intervention. It will also require evidence showing that an agent acted within both its technical permissions and the institution’s approved regulatory logic.

James said his firm is working on “making regulatory reasoning structured, testable and traceable so it can guide human decision-making and support technology-enabled processes.”

The plan’s wider direction suggests that this traceability could become increasingly important. As AI adoption accelerates, banks will need to prove not only that their models work, but that the decisions surrounding those models remain consistent, explainable and defensible.

The emerging QA challenge is therefore broader than model testing alone. It is the testing of the complete decision system: the model, its integrations, the applicable rules, the permitted actions and the evidence showing why each outcome occurred.


NEXT MONTH

REGISTER TODAY – SIMPLY CLICK HERE


Why not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

SIGN UP HERE TODAY


READ MORE


QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS