Banks may need to extend traditional model validation into continuous testing, production monitoring and full-system assurance as artificial intelligence becomes embedded across critical financial workflows, according to the European Central Bank.
Frankfurt-based ECB has made clear that existing governance and model risk frameworks cannot simply be expanded without accounting for the distinct behaviour, dependencies and lifecycle risks introduced by generative and agentic AI.
Pedro Machado, the ECB’s representative to the Supervisory Board of the Single Supervisory Mechanism, recently said more than 85% of large banks under European supervision already use AI in some form, with adoption accelerating as generative and agentic systems move beyond experimentation.
“AI is no longer only about model risk in a narrow sense,” Machado declared in a speech at a KPMG event called RiskTech Conference in Frankfurt earlier this year.
“It increasingly affects governance frameworks, business model evolution and multiple risks – operational risk, conduct risk, compliance risk and strategic risk.”
For quality assurance, software testing and model validation teams, the implications are substantial. AI assurance increasingly needs to cover how models behave under changing conditions, whether outputs can be explained and challenged, how underlying data are governed and whether controls continue to work after deployment.
That moves AI assurance closer to continuous software testing than periodic model approval.
Existing controls not enough?
Banks are not approaching AI without experience. Machado acknowledged that financial institutions have spent decades developing model governance, validation and monitoring controls for credit risk, fraud detection and other regulated applications.
Many banks believe that experience leaves them well placed to manage AI risks. However, the ECB has found that generative and agentic systems introduce qualitative changes that cannot necessarily be addressed by extending conventional controls.
“AI adoption in banking does not start from a blank slate,” Machado said. “The sector has decades of experience with complex models, internal ratings systems and supervisory scrutiny.”
“At the same time, supervisors have also observed that AI introduces qualitative changes that cannot be addressed simply by expanding the existing frameworks.”
The distinction matters for testing teams. Conventional model validation often assesses accuracy, stability and statistical performance at defined approval points.

Generative AI systems can produce different answers to similar prompts, retrieve inappropriate evidence, fabricate convincing conclusions or respond unpredictably when information is incomplete or contradictory.
Agentic systems add another layer of complexity because they may plan tasks, call tools, retrieve information, route decisions between components and take actions with varying degrees of autonomy.
Testing must therefore examine more than the quality of an isolated output. It may need to determine whether the system followed the correct process, used authorised sources and tools, remained within operational boundaries and failed safely when confronted with ambiguous or hostile inputs.
This reflects the wider move towards behavioural AI testing already taking place across the banking sector. Recent work highlighted by Grant Thornton has focused on data quality and safety, behavioural testing, output evaluation, human oversight, continuous monitoring and remediation.
For QA functions, that means building test suites around reasoning consistency, retrieval performance, guardrail effectiveness, refusal behaviour, tool permissions and workflow integrity. It also requires banks to demonstrate that those controls remain effective when models, prompts, data sources and user behaviour change.
Continuous validation
The ECB’s clearest testing message concerns model lifecycle management. Machado warned that machine learning models can evolve as their underlying data change, making monitoring, validation and change management more complicated.
“Consequently, supervisory concerns regarding model management become even more relevant to ensure that model performance can be monitored continuously and to detect drift and unintended effects,” he said.
The ECB also expects “clear escalation and remediation processes” to be in place when models behave unexpectedly.

This creates a direct role for QA and quality engineering teams. Banks need monitoring capable of detecting changes in accuracy, hallucination patterns, retrieval quality, reasoning stability and guardrail performance. Those signals must connect to established defect-management, escalation and remediation processes rather than remain isolated within data science teams.
Testing cannot stop when an AI application enters production. Banks need an evidence trail showing how performance is monitored, which thresholds trigger intervention, who reviews anomalies and how defects or behavioural changes are corrected.
Changes to foundation models, prompts, retrieval sources, APIs and safety controls may also require targeted regression testing. Even where a bank has not altered its own application code, an external model update could change the behaviour of the complete system.
This lifecycle approach echoes the ECB’s wider digital resilience agenda. Frank Elderson, Member of the ECB’s Executive Board and Vice-Chair of its Supervisory Board, has said DORA “provides a regulatory framework that requires banks to foster a culture of continuous improvement in IT and cyber risk management”.
While those remarks concerned operational and cyber resilience, the same principle increasingly applies to AI: controls need to be tested continually and improved as systems and threats evolve.
Explainability becomes testable
Explainability is another area in which the ECB’s expectations go beyond technical documentation.
Machado said banks are increasingly deploying explainability tools but warned that producing an explanation is not sufficient. Banks must establish whether the people accountable for an AI system can meaningfully understand and challenge its decisions.
“It is about ensuring that decision-makers understand what drives model outputs, that risk managers can challenge them, that internal auditors can independently review them and that senior management can take responsibility for how they are used,” he said.
“If a bank cannot explain why an AI model behaves the way it does, in terms that are meaningful for decision-making, then it cannot truly control that model.”
For testing teams, explainability therefore becomes an assurance requirement. Tests may need to establish whether explanations remain consistent with model behaviour, whether decision pathways can be reconstructed and whether auditors can trace an output back to its data, evidence and processing steps.
In customer-facing or decision-making systems, banks may also need to test whether explanations are comprehensible to their intended users, rather than merely accurate in technical terms.
Data quality moves upstream
The ECB is also directing attention away from the model alone and towards the data pipelines that support it.
“AI shifts supervisory attention upstream from models to data,” Machado said. “Data representativeness, data lineage and safeguards against bias are all critical.”

This is particularly important in retail credit, fraud detection and customer segmentation, where historical data may contain structural biases or fail to represent changing customer populations.
The quality challenge is broader for generative AI applications that use retrieval-augmented generation. Testing teams need to check whether source material is current, complete and appropriate; whether access controls prevent the retrieval of restricted information; and whether generated answers are properly grounded in the evidence presented to the system.
The ECB also regards bias as more than a conduct or ethical concern. Machado warned that it can become a prudential issue if it results in the systematic underestimation or mispricing of risk or creates concentration effects.
That makes data-quality testing part of both customer protection and financial risk management.
Assurance before deployment
The ECB expects material AI use cases to be assessed before they are released, with independent challenge from risk management, compliance and internal audit.
Machado said decisions on material applications should be subject to “appropriate pre-implementation assessment and robust risk assessments performed by the second line of defence”.
That requirement strengthens the case for formal AI test gates before production deployment. Banks may need documented evidence covering functional performance, behaviour under adverse conditions, data quality, explainability, security, resilience and human oversight before an application is approved.
Governance remains uneven, however. The ECB has observed fragmented ownership in which responsibility is divided among IT, data science, business lines and control functions without a clear accountability structure.
“AI does not dilute responsibility,” Machado stressed. “If anything, it raises the bar.”
For QA teams, unclear ownership can make defects difficult to escalate and leave gaps between model validation, software testing and business control functions. Banks need clearly assigned responsibility for approving test strategies, accepting residual risks, responding to monitoring alerts and deciding when an AI system must be restricted or withdrawn.
Third-party resilience
Generative AI also introduces dependencies that sit outside banks’ direct control.
Many applications rely on a small group of foundation-model developers, cloud providers and complex subcontracting chains. The models themselves may not be fully transparent to the banks deploying them.
The ECB has linked these dependencies directly to DORA and its wider focus on ICT and third-party risk.
“Generative AI sits at the intersection of technology risk, operational resilience and strategic dependency risk.”
– Pedro Machado
The resulting concerns include cloud concentration, vendor lock-in, confidentiality, security, operational resilience and exit planning. Banks must understand how a failure or material change at a model or cloud provider could affect critical services.
Crucially, contingency arrangements should not exist only on paper. Machado said banks should have “appropriate and tested contingency options for cloud services supporting critical or important functions”.
For quality and resilience teams, this means testing failover arrangements, alternative providers, degraded operating modes, data portability and recovery procedures. Banks may also need to determine whether an AI-supported service can continue safely if a foundation model becomes unavailable, changes unexpectedly or no longer meets regulatory requirements.
The testing burden extends across the supply chain, including external APIs, proprietary platforms, open-source libraries and subcontracted infrastructure.
One supervisory agenda
The ECB’s AI governance, DORA and cyber resilience messages are increasingly converging.
Elderson has warned that advanced AI models are accelerating vulnerability discovery and reducing the time available for banks to test and deploy patches.
He said the growing speed and accessibility of cyber capabilities meant banks needed to prepare “more quickly, more effectively and more consistently across the sector”.

“In musical terms, andante may have previously been good enough, but now we need to move to presto,” Elderson said.
For testing teams, the two sides of the AI challenge are closely connected. Banks must validate the AI systems they deploy while simultaneously protecting their wider technology estate against AI-enabled threats.
DORA’s threat-led penetration testing requirements add to that pressure by asking banks to prove that they can detect, respond to and recover from realistic attacks.
AI behavioural testing applies a similar principle: systems must be evaluated under realistic workflows, ambiguous inputs, missing evidence, malicious prompts and operational stress.
In both cases, a pass-or-fail compliance exercise is unlikely to provide sufficient assurance. Supervisors increasingly expect documented findings, effective escalation, validated remediation and evidence of continuous improvement.
Targeted supervisory scrutiny
AI will remain part of the ECB’s supervisory priorities for 2026–28, particularly under its focus on operational resilience and ICT capabilities.
The regulator intends to continue monitoring AI across the banking sector while applying a more targeted and in-depth approach to generative AI applications. That could pave the way for further supervisory action as the materiality and risks of these systems become clearer.
“Our approach is technology-neutral,” Machado said. “We do not supervise technologies. We supervise how banks apply technologies and ensure good governance, and how this affects their risk profiles.”
“AI must be governed as a core business and risk topic, not as a technology side project.”
– Pedro Machado
For QA and testing leaders, that distinction is important. Banks will not be judged simply on which model, platform or architecture they select.
They will need to show that the resulting system is understood, tested, monitored and controlled throughout its lifecycle.
As AI becomes part of banking’s operational fabric, model risk management, software testing, data assurance, cyber resilience and governance are beginning to merge.
The practical challenge is no longer simply to validate whether an AI model performs well at the point of approval.
Banks must prove that the complete system behaves reliably in real workflows, that failures can be detected and contained, and that controls remain effective as data, models and dependencies change.
Or, as Machado put it: “AI must be governed as a core business and risk topic, not as a technology side project.”
NEXT MONTH



REGISTER TODAY – SIMPLY CLICK HERE
Why not become a QA Financial subscriber?
It’s entirely FREE
* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *
READ MORE
- Goldman puts AI coding to the test
- How to test AI models that banks do not control
- OpenAI, Filigran and SunTec: the latest vendor and product news
- Sygnum: Testing AI is ‘a measurement problem’
- Banks’ ‘code for all’ push raises testing risks
QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS



