Many banks exposed as QA teams battle the ‘looks right’ problem

Khurram Mir

Banks and financial services firms are accelerating their adoption of AI-assisted software development at a pace few engineering organisations could have imagined even two years ago.

From code generation and automated documentation to AI-driven testing and deployment optimisation, the pressure to increase software delivery speed is reshaping how technology teams operate across the industry.

The shift is being driven by a combination of competitive pressure, cost reduction initiatives and the wider race to modernise legacy infrastructure.

Engineering leaders are under mounting pressure to shorten release cycles, expand digital capabilities and integrate AI into development pipelines without slowing innovation. In many organisations, velocity metrics and deployment frequency have become key indicators of progress.

However, beneath the gains in productivity, a growing number of QA and software testing leaders are warning that software quality assurance practices are struggling to evolve at the same pace as AI adoption itself.

The concern is not simply that defects are escaping into production, but that AI-generated software is introducing a new category of risks that often appear technically correct on the surface while masking deeper architectural, integration and logic flaws underneath.

Khurram Javed Mir, founder of Kualitatem and Kualitee, believes this disconnect between AI-driven speed and testing discipline is becoming one of the defining software quality challenges facing financial institutions.

“For the past three years, I’ve watched a pattern repeat itself across software engineering organisations, from early-stage startups to publicly traded software companies,” Mir wrote in a recent analysis.

“The team adopts an AI coding tool. Velocity climbs, leadership celebrates and QA headcount gets quietly frozen or reduced,” he explained.

“Then, six to nine months later, a production incident exposes a logic error that looked perfectly reasonable, passed review and sailed through the rest of the suite.”

The warning comes as financial institutions continue to expand the use of AI-generated code and autonomous development tools inside increasingly complex technology estates. GitHub research cited by Mir found that “AI-assisted developers ship code up to 55% faster.”

However, he cautioned that “output speed and output reliability aren’t the same measurement, and most organizations are only tracking one of them.”

The ‘looks right’ risk

For banks operating interconnected payment systems, authentication layers and customer-facing digital platforms, the issue is not simply defective code, but the emergence of what Mir describes as the “looks right” problem.

“AI code generators are pattern completion engines,” he explained. “They’re extraordinarily good at producing code that resembles correct code.”

Crucially, Mir stressed that the models “are not reasoning about your business logic, your edge cases or the system-level assumptions that a developer, who has since left, baked into your architecture three years ago.”

“The output looks clean because it’s syntactically fluent, not because it’s contextually accurate.”

Khurram Javed Mir

Mir warned that this creates a subtle but increasingly dangerous review dynamic inside engineering organisations, particularly where delivery pressure is intense.

“In code reviews, there’s social pressure to approve,” he said. “When AI-generated code arrives formatted, readable and confident, the bar for pushback rises.”

The risk for banks is that defects increasingly evade traditional review processes before surfacing later in production environments, often within critical business flows.

“The cleaner the output looks on the surface, the more dangerous the blind spot underneath,” Mir warned.

This aligns with broader concerns he raised earlier this year around what he described as the growing erosion of testing discipline across high-velocity software delivery environments.

“The race to accelerate software delivery is exposing a growing fault line for banks and financial services firms,” Mir previously stated. “The risk that speed is outpacing testing discipline, with direct consequences for resilience, customer trust and regulatory exposure.”

QA assumptions break down

A major concern for QA leaders is that many existing testing frameworks were designed around assumptions that no longer hold in AI-assisted development environments.

“Most QA processes were designed around a simple premise: The developer who wrote the code understands it,” Mir wrote. “That premise no longer holds.”

He warned that many organisations are drifting towards what amounts to “circular validation”, where AI-generated code is increasingly tested using AI-generated tests built on similar assumptions and statistical patterns.

“When the same AI, or a similar one, is then used to generate tests for that function, you don’t have quality assurance,” Mir argued. “You have circular validation.”

“The model that produced the code and the model checking it share the same statistical tendencies, the same training blind spots and the same confidence in plausible-looking outputs.”


“The cleaner the output looks on the surface, the more dangerous the blind spot underneath.”

Khurram Javed Mir

For financial services firms operating under DORA, operational resilience frameworks and increasing regulatory scrutiny around software governance, the implications are significant.

Failures may no longer originate from obvious coding mistakes, but from hidden assumptions, undocumented dependencies and integration-level weaknesses that only emerge under real-world conditions.

“This is the new technical debt,” Mir warned. “It doesn’t show up in your sprint metrics or trigger alerts. It accumulates until a production incident forces the conversation nobody wanted to have at scale.”

From execution to interpretation

Rather than reducing QA investment as AI accelerates delivery, Mir argued the opposite is required.

“The instinct to reduce QA investment as AI output scales is almost exactly backwards,” he wrote. “More AI-generated code means more output that requires interpretive review.”

The organisations navigating the shift most effectively are separating code generation from test strategy, according to Mir.

“AI handles the execution layer. Humans own the test strategy,” he stated.

For QA teams inside banks, this increasingly means moving away from purely functional testing towards interpretive and behavioural validation.

“What was this code supposed to do?” Mir asked. “What assumption is it making about the data it receives?”

He added that forward-looking engineering organisations are increasingly prioritising behavioural and integration testing over pure unit test volume.

“Unit tests confirm that individual functions execute,” he explained. “Integration tests confirm that systems behave as intended under real conditions.”

That distinction is becoming increasingly critical for financial institutions, where AI-generated code may pass isolated functional checks while introducing hidden risks across broader transaction flows and interconnected systems.

“Automation coverage is often treated as a vanity metric,” he said in earlier remarks on testing discipline. “Teams attempt to automate everything, producing suites that become slow and brittle.”

Mir also reiterated concerns around organisations focusing too heavily on automation coverage as a headline metric.

Instead, he advocates a risk-based testing approach prioritising “payment systems, authentication flows, compliance processes and service integrations.”

‘Capital protection’ strategy

As banks continue scaling AI-assisted development, Mir believes QA is rapidly becoming a frontline resilience function rather than a background engineering process.

“The teams that will earn trust in the AI era aren’t the ones shipping fastest,” he concluded. “They’re the ones whose AI-assisted output actually holds up when customers use it.”

He warned that organisations ignoring the widening gap between AI adoption and testing maturity are exposing themselves to systemic operational risk.

“Does your testing discipline scale as quickly as your AI adoption does?” Mir asked. “If you can’t answer that confidently, the gap between those two curves is exactly where your next crisis is forming.”

Ultimately, he argued that financial institutions must rethink how they position software quality internally.

“Speed without discipline almost always introduces operational risk that eventually becomes a business constraint,” Mir stated.

“They will have the most success in preventing this issue if they see testing discipline as a capital protection strategy rather than an engineering preference.”


16 SEPTEMBER IN LONDON

REGISTER TODAY – SIMPLY CLICK HERE


Why not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

SIGN UP HERE TODAY


REGULATION & COMPLIANCE

Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.


READ MORE


WATCH NOW


QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS