Bank of England raises the testing bar for frontier AI

Andrew Bailey

Bank of England Governor Andrew Bailey is putting stress testing, penetration testing and pre-deployment model assessment at the centre of the response to AI-driven cyber threats, while prompting a wider debate about how regulators demonstrate the effectiveness of their own controls.

Bailey intervened after Daily Mail columnist Connor Axiotes described the Bank’s cyber defences as “unsophisticated”.

In an open letter sent to the newspaper and published by the Bank on July 23, the Governor rejected the assertion as “totally wrong, unfounded and dangerous”.

“For obvious security reasons we do not discuss our defences in detail,” Bailey wrote.

However, Bailey assured readers that the Bank has “the right investment, expertise and capabilities” and works closely with specialist partners, including the National Cyber Security Centre.


“The Bank calls for stronger international coordination around testing frontier AI models.”

– Andrew Bailey

The exchange exposes a familiar challenge for quality assurance and cyber resilience teams: how can an institution provide credible evidence that its systems have been rigorously tested without revealing information that could help potential attackers?

Bailey shifted the focus from the Bank’s own controls to the wider risks that frontier AI poses to financial institutions and their customers.

“Frontier AI may make cyber-attacks faster and easier to perpetrate, outages more disruptive, and scams by criminals more convincing,” he warned.

The Bank has repeatedly told firms to strengthen their ability to identify attacks, address vulnerabilities and restore services following disruption.

The Bank of England in the heart of the City of London

Bailey said banks must “strengthen their detection efforts and responses, patch vulnerabilities faster, and be able to recover when things do go wrong.”

Crucially for testing teams, these expectations cannot be met through policies or assurances alone.

“As a regulator we require that banks prove to us that they can do this through stress tests and penetration testing,” Bailey wrote.

That statement puts demonstrable testing evidence at the heart of cyber resilience. Financial institutions must be able to show that vulnerabilities have been identified, adverse scenarios have been exercised and recovery arrangements work under realistic conditions.

It also raises a question about the regulator itself. The Bank may have legitimate reasons for withholding sensitive security information, but its position invites scrutiny over how its own testing evidence is independently assessed.

‘Trust us’

Abdelhamid Taha, a member of the secretariat of the All-Party Parliamentary Group on Investment Fraud and Fairer Financial Services, highlighted that tension in a LinkedIn post, in which he responded to Bailey’s letter.

Taha said he was not questioning whether the Bank’s cyber defences were effective and acknowledged that it could not reasonably publish the details of its security architecture in a newspaper.

Abdelhamid Taha

“What caught my attention is the shape of the answer,” he outlined, arguing that he had encountered the same pattern while examining FCA case files involving financial harm.

Taha characterised that pattern as: “We can’t show you the evidence, but trust our judgment that we’ve got this handled.”

He stressed that such responses are not necessarily dishonest or intended to conceal misconduct. His concern is that authoritative assurances may be accepted without the underlying evidence being available for wider examination.

“Not lies. Not cover-ups. Just: ‘trust us, [but] we can’t show our working,’” Taha argued.

That distinction is important, he stressed, as cybersecurity secrecy does not eliminate the need for evidence; it determines who can review that evidence and under what conditions.

Sensitive test results can be examined confidentially by authorised oversight bodies, independent assessors and internal assurance teams. The central issue is whether statements about resilience can be traced to repeatable testing, documented findings, remediation records and verified recovery performance.

Testing the regulator’s AI

Taha connected the exchange to the FCA’s planned use of agentic AI in financial supervision. He referred to an Agentic Supervisory Model that would monitor the financial sector and help identify emerging problems more rapidly than conventional supervisory processes.

“If ‘trust us, we can’t show our working’ is the pattern baked into that training data, the AI won’t learn to question that pattern, it’ll learn to reproduce it,” Taha argued.

He warned that such a system could offer confident assurances because that is what its training material taught it “sounding credible” looks like.

Taha’s claim about the prospective system’s training material is an argument rather than a disclosed feature of the FCA’s technical design. Nevertheless, it identifies a credible QA risk that applies regardless of the precise datasets eventually selected.

A supervisory AI would need to distinguish between substantiated assurance and authoritative language unsupported by accessible evidence. Testing should establish how the system responds when a bank, regulator or senior executive says controls are effective but cannot disclose the supporting material.

The system should be capable of recognising legitimate confidentiality constraints while still requesting appropriate evidence, recording uncertainty and escalating claims that cannot be independently verified.

That would require testing for institutional bias, automation bias, explainability and false reassurance. Adversarial scenarios could also examine whether the status or confidence of a source influences the model more strongly than the evidence available to support its claims.

Testing before deployment

Bailey also used his letter to call for stronger international cooperation on the assessment of advanced AI systems.

“It is also why the Bank has called for stronger international coordination around testing frontier AI models before wider deployment,” he wrote.

The Governor added: “No country can seal itself off from these risks.”

Bailey pointed to the UK’s AI Security Institute and the NCSC, with the Bank working alongside domestic and international partners to reinforce the cyber resilience of supervised financial institutions.

“The issues of AI and cyber-attack are rightly being scrutinised,” Bailey concluded, assuring readers that the Bank and other UK authorities are “focused on them and taking action.”

For banks, the letter reinforces the expectation that resilience should be proven through testing rather than asserted through policy. It also establishes a demanding benchmark for regulators as they introduce AI into their own operations.

If financial institutions must prove their resilience through stress tests and penetration testing, confidence in regulatory AI will similarly depend on evidence that the technology has been tested to challenge unsupported assertions, not merely reproduce them convincingly.


NEXT MONTH

REGISTER TODAY – SIMPLY CLICK HERE


Why not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

SIGN UP HERE TODAY


READ MORE


QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS