
Germany’s financial services regulator has begun monitoring artificial intelligence across the country’s banks and insurers, sharpening the supervisory consequences of its earlier guidance on AI testing and operational resilience.
BaFin has told financial firms to apply established controls, including unit, integration testing, and adversarial testing, to most of its integrated AI software.
The watchdog’s expanded authority increases pressure on banks to test and demonstrate that customer-facing and decision-making systems remain transparent, secure and free from discrimination.
The regulator’s new powers took effect on July 29 following legislation that allows the Frankfurt-based body to impose fines.
BaFin will initially monitor compliance with transparency requirements, including whether customers are informed when they interact with chatbots. It will also test and examine the use of AI in sensitive financial decisions and check that firms are not engaging in prohibited practices involving personal data.
High-risk AI software
In addition, from December 2027, BaFin’s remit will extend to high-risk AI systems. Within financial services, this is expected to include technology used to assess consumers’ creditworthiness and produce credit scores.
The move puts AI testing more firmly within Germany’s regulatory landscape. Banks will need effective controls for assessing whether systems remain accurate, explainable and non-discriminatory as underlying data, models and connected applications change.
“People have to be able to trust that their fundamental rights will be protected when AI is used,” stressed BaFin president Mark Branson.

“BaFin will ensure, for example, that everyone has fair access to financial services and that no one is discriminated against as a result of AI,” he added.
BaFin will focus on AI used in activities that require regulatory authorisation, including banking and insurance services, according to Branson’s colleague, Jens Obermöller, director-General for Cyber Risks and Technology at BaFin.
Rather than inspecting every model operated by every institution, the watchdog plans to review samples of applications used in particularly significant areas, he stated.
“For example, we will review a sample of the AI applications that are used by many financial entities in particularly relevant areas,” explained Obermöller.
“But what we do not do is monitor every single AI system in every financial entity. What the act mandates is not supervision; it is monitoring.”
However, Obermöller did stress that, “for this purpose, regulatory sandboxes and testing in real world conditions are planned. In these secure environments, banks can test new applications.”
That distinction does not remove the need for banks to maintain reliable information about how their systems have been tested and governed.
If an application is selected for review, an institution may need to explain how the technology was validated, how potential discrimination is detected and how its performance is monitored after deployment.
“Regulatory sandboxes and testing in real world conditions are planned. In these secure environments, banks can test new applications.”
– Jens Obermöller
Testing creditworthiness systems is likely to prove particularly demanding. Financial institutions must test, monitor and determine whether model outputs remain accurate across different groups and whether changes to data, models or integrations introduce unfair outcomes.
Banks also need to test and supervise chatbots to ensure they work consistently across customer journeys, devices, languages and software releases.
A disclosure policy offers limited protection if a defect, configuration problem or third-party update prevents the required information from reaching customers.
Resilience framework
BaFin’s new monitoring powers build on guidance issued earlier this year that placed AI within the operational resilience agenda for banks and insurers.
The regulator said AI should be managed within the same ICT risk framework as other important technology assets under the EU’s Digital Operational Resilience Act.
Its security and resilience should therefore be maintained across the full lifecycle, from data acquisition and development to deployment, operation and eventual retirement.

Marina Marusenko, a risk manager at Dutch bank ING, an active player in the German market, summarised the regulator’s position by saying: “BaFin treated AI as an ICT risk issue, not an innovation topic.”
That approach has direct consequences for QA teams. AI assurance can no longer be treated solely as a data-science exercise or a final check conducted immediately before deployment. It becomes part of the institution’s wider framework for identifying, preventing, detecting and recovering from ICT disruption.
BaFin’s guidance stressed that established software-engineering disciplines, including unit testing, integration testing and source-code reviews, remained important when validating AI systems. The extent of testing should reflect the criticality of the business functions supported by the technology.
Its guidance also placed lifecycle management at the centre of AI assurance, covering development, deployment, monitoring, change management and decommissioning.
“A one-time approval at go-live was not sufficient,” Marusenko wrote on LinkedIn, pointing to the need for continued oversight of model drift, data quality and version control.
This is particularly relevant when banks use generative AI supplied by external providers. Models can change after their original approval, sometimes without the financial institution controlling, or immediately knowing about, the update.
Regression testing, output monitoring and model-version controls consequently become essential to determining whether a previously approved system continues to behave as expected.
Testing moves beyond model accuracy
BaFin’s earlier guidance also pushed financial institutions beyond conventional functional testing and towards adversarial resilience assessments.

The regulator highlighted attack simulations involving data poisoning and evasion, alongside penetration testing intended to identify AI-specific vulnerabilities.
It warned that models could be compromised during training or manipulated after deployment, potentially causing a financial institution to make incorrect decisions.
For banks, this means testing must cover more than whether an AI system produces accurate results under ordinary conditions. QA and security teams also need to assess how it behaves when its data, prompts, interfaces or underlying infrastructure are deliberately manipulated.
The regulatory focus on fundamental rights adds another testing layer. Systems affecting access to financial services require controls capable of detecting whether particular groups receive systematically different outcomes and whether those differences can be justified.
Explainability must also work in practice. If a bank cannot trace or meaningfully explain how an AI-supported decision was reached, it may struggle to investigate defects, address customer complaints or demonstrate fair treatment to BaFin.
Third-party models
BaFin has repeatedly drawn attention to financial institutions’ dependence on cloud platforms and a relatively small number of AI providers.
Its resilience guidance urged firms to assess vendor lock-in, secure meaningful audit rights and maintain credible strategies for transferring models and training data or exiting a provider altogether.
Testing those exit plans will be important. A written portability commitment does not establish that a bank can move an AI service without losing functionality, historical data, monitoring capabilities or control over customer outcomes.
External application programming interfaces create an additional visibility problem. Software originally classified as a conventional application can become an AI system when a team connects it to an external model, potentially leaving governance and testing teams unaware of the resulting exposure.
Identifying where AI is actually used across the technology estate is therefore a prerequisite for testing it. Without an accurate inventory, banks cannot reliably determine which applications require AI-specific validation, monitoring or resilience controls.
NEXT MONTH



REGISTER TODAY – SIMPLY CLICK HERE
Why not become a QA Financial subscriber?
It’s entirely FREE
* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *
REGULATION & COMPLIANCE
Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.
READ MORE
- Goldman puts AI coding to the test
- How to test AI models that banks do not control
- OpenAI, Filigran and SunTec: the latest vendor and product news
- Sygnum: Testing AI is ‘a measurement problem’
- Banks’ ‘code for all’ push raises testing risks
WATCH NOW


QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS



