EU AI rules shift compliance to testing

The European Union has spent much of the past year redefining software testing as a cornerstone of AI governance rather than simply a quality assurance exercise.

BaFin has positioned generative AI as an operational resilience issue under the Digital Operational Resilience Act (DORA), the European Commission’s guidance on high-risk AI systems has expanded expectations around lifecycle validation and evidence, and supervisors across the region have increasingly asked financial institutions to demonstrate not only that AI systems work, but that they remain governed, monitored and under control after deployment.

The latest development continues that trajectory. Earlier this month, the European Commission published its voluntary Code of Practice on AI-generated content, providing providers and deployers with practical guidance on how to meet the transparency obligations of the EU AI Act ahead of their application from 2 August.

Developed by independent experts and intended to help organisations implement the Act’s transparency provisions consistently, the Code sets out recommended approaches for identifying and disclosing AI-generated and AI-manipulated content through machine-readable metadata, provenance information, watermarking and other technical mechanisms.

At first glance, the document appears aimed primarily at AI developers, content platforms and providers of general-purpose AI models. Much of the early discussion has centred on deepfakes, disclosure labels and watermarking.

For banks and financial institutions, however, the implications extend much further.

As organisations embed generative AI into customer communications, document processing, software engineering, internal copilots, compliance workflows and operational processes, every transparency label, watermark and provenance record becomes another control that must survive software updates, API integrations, document conversions and increasingly complex enterprise technology stacks.

That makes this far more than a transparency initiative. It is another indication that European regulators increasingly expect AI governance to be demonstrated through testing.

Just as BaFin argued that AI belongs within the same ICT risk management framework as every other critical technology asset, and the Commission’s recent guidance on high-risk AI systems pointed towards continuous validation and evidence rather than one-off approval, the new Code effectively creates another category of assurance work for software testing and quality engineering teams.

The challenge is no longer simply adding an AI label to generated content. It is proving that the label remains accurate, intact and auditable wherever that content travels.

From transparency to testing

The Code is built around a relatively straightforward objective: ensuring AI-generated content can be identified and its origin understood.

Achieving that objective inside a modern bank is considerably more complex. A customer letter drafted using a generative AI assistant may be edited by several employees, converted into PDF, uploaded into an enterprise content management platform, sent through email, archived for regulatory purposes and ultimately viewed inside a customer portal.

Each stage represents another opportunity for transparency information to disappear. Metadata can be stripped during document conversion. APIs may fail to preserve provenance information between systems.

Content copied into another application may lose its disclosure markers altogether. Third-party software updates may inadvertently remove or overwrite transparency controls that previously functioned correctly.

For software testing teams, these become entirely new validation scenarios. Does every AI-generated output receive the appropriate disclosure? Does provenance survive every file conversion?

Moreover, can transparency information withstand routine software updates? What happens when multiple AI systems successively modify the same document? Can labels be removed accidentally, or deliberately? These are no longer simply functional testing questions. They are governance questions.

Testing as regulatory evidence

That reflects a broader shift already emerging across European AI regulation. When the Commission published its guidance on high-risk AI systems earlier this year, the document appeared at first to focus on legal classification.

Yet for financial institutions, its more significant consequence was the expectation that organisations document, monitor and govern AI systems throughout their lifecycle.

BaFin’s HQ

Increasingly, regulators are asking organisations not merely whether an AI system performs correctly, but whether they can demonstrate how risks are identified, how human oversight is maintained, how changes are controlled and how decisions are monitored over time.

Testing therefore becomes much more than a release activity. It becomes part of the evidence supporting regulatory compliance.

That same theme runs through BaFin’s AI guidance, which deliberately avoided treating artificial intelligence as a standalone innovation topic.

Instead, the German regulator positioned AI alongside every other critical ICT asset, arguing that AI systems should be governed using the same operational resilience principles applied to the wider technology estate.

The guidance also made another important point. Development, deployment and go-live approval are no longer enough.

Instead, AI systems require continuous monitoring, lifecycle management, change control and ongoing validation as models, data and business processes evolve. For quality engineering teams, that represents a significant expansion of responsibility.

Continuous assurance replaces one-off approval

Traditional software projects often culminated in a successful production release. AI governance increasingly starts there.

As banks accelerate their use of generative AI, transparency controls will need to become part of continuous integration and continuous delivery pipelines alongside existing security, accessibility and operational resilience checks.

Regression suites may need to verify that provenance metadata survives every software release. Integration testing will need to confirm that transparency information remains intact as content moves between cloud platforms, internal applications and third-party services.

Automated monitoring may be required to detect when software updates unexpectedly break disclosure mechanisms.

Audit logs may need to demonstrate exactly when AI-generated content was created, modified and distributed. In effect, transparency becomes another compliance control requiring continuous validation rather than periodic review.

The pattern will feel familiar to many banking technology teams. Accessibility testing evolved from a specialist exercise into a continuous engineering discipline as regulation matured.

Cyber resilience followed a similar path under DORA, where testing increasingly provides evidence that critical services remain operational under realistic failure scenarios. AI transparency appears to be following the same trajectory.

Beyond functional testing

The testing challenge extends beyond confirming that labels simply exist. Engineering teams will increasingly need to understand how transparency controls behave when systems fail, when metadata becomes corrupted or when content passes through applications that were never designed with AI provenance in mind.

Negative testing therefore becomes increasingly important. Can disclosure labels survive copy-and-paste operations? What happens when documents are compressed, reformatted or exported into different file types?

Also, can provenance metadata be manipulated by malicious actors? Are APIs consistently preserving machine-readable information across multiple platforms?

Detection technologies require scrutiny too. As organisations deploy systems capable of recognising AI-generated or AI-manipulated content, those tools themselves become subject to testing.

False positives, false negatives, multilingual performance, robustness against edited content and reliability across different document formats all become measurable quality characteristics.

For financial institutions, these are not merely technical questions. Incorrectly identifying genuine customer communications as AI-generated could disrupt customer journeys and create compliance issues.

Failing to identify synthetic content may expose institutions to fraud, operational risk and reputational damage. Quality engineering therefore moves closer to fraud prevention, cyber resilience and enterprise risk management.

Another step in Europe’s AI assurance agenda

Viewed in isolation, the Commission’s Code of Practice is a transparency document. Viewed alongside BaFin’s operational resilience guidance, the Commission’s high-risk AI guidelines and DORA’s growing emphasis on continuous technology assurance, it represents another milestone in Europe’s broader AI governance strategy.

Across each of these initiatives, one theme continues to emerge. Regulators are becoming less interested in whether organisations say controls exist, and more interested in whether they can prove those controls continue working as systems evolve.

That inevitably shifts software testing into a more strategic role. Quality engineering is no longer simply responsible for identifying defects before production.

Increasingly, it is expected to provide continuous, repeatable and auditable evidence that AI systems remain governed, transparent, resilient and compliant throughout their operational lifecycle.

For banks, that means AI governance is unlikely to be judged solely by policies, committees or model documentation. It will increasingly be measured by something far more familiar to testing teams: whether organisations can demonstrate, release after release, that their controls continue working under real-world conditions.


THIS SEPTEMBER IN LONDON

REGISTER TODAY – SIMPLY CLICK HERE


Why not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

SIGN UP HERE TODAY


READ MORE


QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS