AI quality failures become governance failures for banks, warns Testlio CPTO

Indiana, U.S.-based Darin Brown

As banks and financial institutions accelerate the deployment of AI-powered customer service, recommendation engines and digital assistants, software testing teams are being pushed into unfamiliar territory where traditional QA metrics no longer capture the failures that matter most.

According to Darin Brown, Chief Product and Technology Officer at Testlio, the challenge facing financial services firms is no longer simply whether an AI system passes a test suite, but whether it behaves responsibly under unpredictable real-world conditions.

“There is a real difference between a product passing a test suite and actually being a good experience,” Brown said. “That gap is where quality lives now, and in financial services it is widening.”

For banks operating under growing regulatory scrutiny around operational resilience, AI governance and customer protection, that distinction is becoming increasingly important.

An AI system may technically perform correctly while still damaging trust, creating compliance exposure or driving customer churn.

Brown pointed to scenarios where AI systems produce technically acceptable outcomes that nonetheless undermine the customer experience.

“The payment is technically processing, but the latency is just long enough to make the customer doubt it went through,” he said. “The chatbot is technically accurate about the overdraft policy, but it answers in a tone that reads as dismissive to a customer in real financial distress.”

He added that recommendation engines can also fail in more subtle ways. “The recommendation is technically aligned with the customer’s stated preferences, but it ignores context a human would have caught from two messages earlier in the conversation.”


“There is a difference between a product passing a test and being a good experience. That gap is where quality lives now, and in financial services it is widening.”

– Darin Brown

None of those failures, Brown argued, are likely to appear in conventional evaluation suites. “All of them show up in churn.”

That challenge is becoming more acute as financial institutions race to deploy AI-powered digital banking experiences.

Brown referenced research from a recent report called State of Digital Banking showing that “66% of consumers demand a seamless experience across all their digital banking platforms,” warning that “AI alone is not going to clear it.”

He also warned that generative AI systems can create particular risks inside regulated environments because “AI is very good at sounding authoritative when it is wrong.” In banking environments, he said, “that authority bias is dangerous.”

For QA and software testing teams, Brown believes the definition of quality itself is changing. “The definition of quality has moved from ‘does it work’ to ‘does it behave responsibly under conditions no one fully anticipated.’”

AI hallucinations now a governance problem

Brown argued that many banks are still approaching hallucinations and misinformation as isolated testing defects rather than governance issues that cut across product, compliance and engineering functions.

“Hallucination is the most visible instance of the broader problem,” he said. “A wrong answer from a chatbot is a customer experience issue at best. At worst, it is a regulatory violation, bad financial advice given under the bank’s brand, or a privacy breach the customer discovers before the institution does.”

But Brown pushed back strongly on the idea that hallucinations can be solved through automated testing alone.

“The framing I would push back on is that this is a testing problem,” he said. “It is a governance problem that shows up in testing.”

According to Brown, determining whether an AI-generated response about a banking product or loan policy is acceptable requires policy decisions that sit beyond the authority of most traditional QA functions.

“Determining whether a chatbot’s answer about a loan policy is acceptable is not a test case,” he said. “It is a policy decision about what the model is allowed to assert, in which contexts, and with what confidence.”


“Hallucinations is a governance problem that shows up in testing.”

– Darin Brown

That challenge is becoming more visible as banks experiment with automated AI evaluation systems where one model judges the outputs of another.

Brown acknowledged those approaches can help with “format compliance and policy adherence at scale,” but warned that they are insufficient for the kinds of failures that create real damage inside financial institutions.

“The evaluator inherits the same blind spots as the system under test,” he said. “Both models were trained on similar data, both reflect similar assumptions about what a reasonable answer looks like, and neither has standing to judge whether a specific answer is appropriate for a specific customer in a specific regulatory context.”

Brown said his firm’s own approach combines proprietary AI tooling with human reviewers capable of challenging assumptions and identifying failures internal teams may overlook.

“The human-in-the-loop layer matters most where domain judgment and outside perspective are non-negotiable,” he stressed.

He warned that internal engineering and QA teams often develop blind spots over time because they become too familiar with how a system is intended to behave.

“They know what the AI is supposed to do and instinctively avoid the questions that break it,” Brown said.

“Outside experts come in without those reflexes, push the boundary conditions, and surface the failures that automated suites and internal reviewers consistently miss.”

Structural weaknesses in AI testing models

Brown believes the biggest weakness in how banks currently test AI systems is not tooling, but organisational structure.

“The biggest gap is structural,” he said. “Most banks have an AI quality function that is owned by a single internal team and validated primarily through automated suites.”

That model, he argued, “almost guarantees the failures that matter will be missed.”

The commercial risks are significant. Brown cited research from Acquire BPO indicating that “70% of customers will switch providers after a single bad AI interaction.”

“In financial services, where trust compounds over years and erodes in hours, you do not need most customers to leave,” he said.

“You need a small number of high-profile failures to reach the regulator, the press, or a competitor’s marketing team.”

Brown also referenced concerns already raised by the Consumer Financial Protection Bureau around poorly deployed banking chatbots and the compliance risks they create.

For fintechs and neobanks, the pressure may be even greater. “A single viral incident represents a much larger share of the brand than it would for an incumbent,” Brown said.

He repeatedly stressed that banks should not view AI quality as solely an engineering or QA responsibility.

“The gap is not a tooling gap,” Brown said. “It is an operating model gap.”

Rise of cross-functional AI quality

As firms prepare for growing regulatory obligations under frameworks such as Digital Operational Resilience Act, Brown believes QA teams will need to operate much more closely with product, engineering, compliance and design functions.

“DORA is forcing financial firms to articulate what their systems are permitted to do, how they detect failure, and how they respond inside a defined window,” he explained.

Brown argued that engineering, product, design and QA teams must each own distinct parts of the AI quality process.

“Engineering owns model and infrastructure integrity,” he said. “Product owns decision boundaries.” Meanwhile, “Design and QA own the user experience.”

Importantly, Brown believes the tension between those groups is necessary rather than problematic. “The tension between those three is the system working,” he shared.


“DORA is forcing financial firms to articulate what their systems are permitted to do.”

– Darin Brown

Brown warned that failure emerges when a single function dominates AI decision-making internally and conflicting priorities disappear from the validation process.

He also argued that adversarial testing will become increasingly important for banks deploying AI systems into customer-facing channels.

“Guardrails are not standardized,” Brown continued. “They have to be adapted to a bank’s specific business context.”

Even sophisticated safeguards remain vulnerable, he warned, to “poetic jailbreaks and prompt injections embedded in customer-submitted content.”

For QA and testing leaders inside financial institutions, Brown believes the focus now needs to shift from launch readiness toward continuous operational oversight.

“The institutions that get this right will not be the ones with the most sophisticated AI,” he stated.

“They will be the ones with the clearest model of who owns which question, the discipline to put genuine outside judgment into the validation loop before a feature reaches a customer,” Brown concluded. “And the operational maturity to treat AI quality as a continuous obligation rather than a launch checklist.”


WHY not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

REGISTER HERE TODAY


READ MORE


WATCH NOW


QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS