ECB pushes banks to behavioural AI testing

The ECB's headquarters in Frankfurt, Germany
The ECB's headquarters in Frankfurt, Germany

Banks are being pushed towards a new era of AI assurance in which traditional model validation is no longer enough, with behavioural testing, output evaluation and continuous monitoring becoming central to how financial firms prove that generative and agentic AI systems can be trusted.

The shift is highlighted in a recent Grant Thornton analysis on responsible AI in banking and AI system validation, which argues that banks need to move beyond conventional model risk controls as AI systems become more embedded in critical workflows.

The message is significant: AI assurance is starting to look less like periodic model approval and more like continuous software testing.

GT’s validation framework focuses on data quality and safety, behavioural testing, output evaluation, human oversight, continuous monitoring and remediation. The firm wrote that behavioural testing is needed to assess whether generative AI and agentic AI systems “behave with appropriate discipline and control in practice”.

That includes how consistently an AI system reasons, how reliably it grounds outputs in available evidence, how it responds when information is missing or contradictory, and whether guardrails remain effective over time.

Testing challenge

For agentic AI systems, the testing challenge becomes even broader. These systems may plan tasks, retrieve information, route decisions between components, use tools or APIs and act with varying levels of autonomy. That introduces risks around planning integrity, workflow coherence, tool-use safety, state integrity, retrieval behaviour, auditability and guardrail failures.

Lukas Majer

In a LinkedIn post following Grant Thornton’s Responsible AI in Banking conference, Lukas Majer, Head of Quantitative Risk at Grant Thornton in Madrid, wrote that the firm’s discussions had involved the European Central Bank, BBVA, Santander, Commerzbank, Lloyds Banking Group and Grant Thornton.

He disclosed that they covered “behavioural testing, output evaluation, human oversight, continuous monitoring and remediation mechanisms to help organisations build trustworthy AI capabilities.”

The focus on behaviour is important because many AI failures do not resemble traditional software defects. Generative AI systems may produce fluent but unsupported answers, shift reasoning patterns between similar prompts, fail to refuse unsafe or incomplete requests, retrieve the wrong evidence or present fabricated conclusions with confidence.

Agentic systems add further complexity because they may take multi-step actions. In that context, banks do not only need to test whether an output is correct. They need to test how the system reached that output, whether it used the right sources, whether it stayed within permitted boundaries and whether its behaviour remains stable over time.

That is where behavioural testing becomes directly relevant for QA teams. Grant Thornton said behavioural testing focuses on whether an AI system “behaves safely, predictably, and consistently across different conditions, rather than assessing the quality of any single output in isolation”.

The firm also argued that continuous monitoring should track drift, hallucination patterns, retrieval failures, planning instability and other behavioural changes over time.

ECB direction

This aligns closely with the direction of European supervisory expectations. The European Central Bank has already made clear that AI is becoming a more prominent supervisory concern, particularly as banks move from experimentation towards operational deployment.

In February, the ECB said it would continue to monitor AI use across banks while taking a more targeted and in-depth approach to generative AI applications.

The central bank said its objective is not to supervise technologies themselves, but to supervise how banks apply technologies, how they govern them and how those technologies affect their risk profiles.

That framing creates a clear opening for QA and testing leaders. If supervisors increasingly care about how AI behaves in production, banks will need evidence that behaviour has been tested, monitored and remediated.

Frank Elderson

The same logic is visible in the ECB’s cyber resilience agenda. Frank Elderson, Member of the Executive Board of the ECB and Vice-Chair of the Supervisory Board of the ECB, recently warned that AI is changing the speed and scale of cyber risk.

“The direction of travel is unmistakable: the speed, scale and accessibility of advanced cyber capabilities are increasing, and the time available to defenders is shrinking,” Elderson said.

“Banks therefore need to prepare more quickly, more effectively and more consistently across the sector,” he added.

“In musical terms, andante may have previously been good enough, but now we need to move to presto.”

While those remarks focused on cyber risk and DORA testing, the same principle applies to AI assurance. Static reviews and one-off approvals are unlikely to be sufficient for systems whose behaviour can vary with prompts, data, retrieval context, model updates and user interaction.

DORA has already raised expectations around continuous improvement, ICT testing and resilience evidence. Elderson said DORA “provides a regulatory framework that requires banks to foster a culture of continuous improvement in IT and cyber risk management”.

He also said DORA “gave supervisors the task of testing whether a financial institution can detect, respond to and recover from sophisticated attacks that mirror real-world threats, thereby providing a more systemic and enforceable framework for resilience.”

For software testing teams, that emphasis on realistic testing is increasingly relevant to AI. Behavioural testing of AI systems is, in effect, a way of asking similar questions: does the system behave as expected under realistic conditions, does it fail safely, can the firm detect when behaviour changes, and is there a documented remediation loop?

Threat-led penetration testing

The ECB’s TIBER-EU approach to threat-led penetration testing also reinforces the same direction of travel. TIBER-EU is designed to expose how systems and teams behave under realistic attack conditions rather than simply produce a pass-or-fail result.

AI validation is moving in a similar direction. The question is no longer only whether a model performs well against a benchmark. The question is whether the full AI system behaves reliably across real workflows, changing inputs, ambiguous information, missing evidence and operational stress.

That has major implications. Testing teams inside banks are likely to become more involved in AI assurance, not less, as model risk management, software testing, operational resilience and governance begin to converge.

The practical work will include designing behavioural test suites, validating prompt and retrieval performance, testing refusals and guardrails, checking tool-use boundaries, monitoring drift, documenting defects, validating remediation and producing evidence that supervisors, audit teams and senior management can understand.

Grant Thornton’s analysis makes clear that AI validation requires a lifecycle approach. Even with strong controls, generative and agentic systems require continuous remediation because issues can arise at any stage.

For banks, this means AI testing cannot sit only at the model approval stage. It has to extend into development, deployment, monitoring and change management.

That is the key QA takeaway: As AI becomes part of banking infrastructure, behavioural testing is emerging as one of the most important tools for proving that these systems are safe, controlled and fit for purpose. Or, put more simply: banks are no longer just validating models. They are testing behaviour.


16 SEPTEMBER IN LONDON

REGISTER TODAY – SIMPLY CLICK HERE


Why not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

SIGN UP HERE TODAY


REGULATION & COMPLIANCE

Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.


READ MORE


WATCH NOW


QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS