Do AI agents bring ‘error inheritance’ risk to banks?

Colorado-based Benjamin Walker

As banks deploy AI assistants across customer-service operations, software testing teams are facing a new quality challenge: ensuring that errors do not spread unchecked through increasingly complex AI-driven workflows.

The issue has come into sharper focus following Bank of America’s expansion of EricaAssist, its human-assisted AI system used by more than 18,000 customer-service representatives.

The technology can summarise why a customer is calling, retrieve relevant information and recommend next steps in under three seconds. According to the bank, the system has also reduced average call times by nearly one minute.

While the speed gains are significant, quality engineers face a different question. As AI-generated transcripts and summaries become inputs for other systems, a single recognition error can potentially be carried through multiple stages of a banking workflow, influencing decisions, analytics and future AI outputs.

Denver, Colorado-based Benjamin Walker, CEO of transcription company Ditto Transcripts, believes this creates what he describes as an “error inheritance” risk.

“Three seconds is impressive, but banking decisions can follow a customer for years,” Walker pointed out.

“If a system mishears a disputed amount, misses the word ‘not’ or assigns a statement to the wrong speaker, that mistake may appear in a summary that looks polished and authoritative,” he explained.

Walker added that “the danger is that nobody reopens the recording because the summary reads so confidently.”

Hallucinations

The challenge extends well beyond verifying the performance of an individual AI model. Banks increasingly need to validate the entire decision chain, from speech recognition and transcript generation through AI summaries, recommendations and the downstream systems that consume those outputs.

An isolated transcription mistake may no longer remain an isolated defect if it propagates across customer records, analytics platforms, complaint investigations and future AI training data.

The growing reliance on AI-generated information comes as financial institutions continue to express concerns about the reliability of AI outputs.

This year’s Global AI in Financial Services Report from Cambridge Judge Business School found that hallucinations and unreliable AI outputs concerned 70% of surveyed financial institutions. Data availability and quality were also cited as barriers by 40% of industry respondents.

Similarly, an ACA Group survey of more than 200 US financial-services firms found that 84% were using AI in some capacity.


“Banks should decide how much human review is required according to how the record will be used.”

– Benjamin Walker

However, active deployment averaged below 20% across compliance functions and approximately 5% across operations. AI-generated errors, regulatory scrutiny and the need for clear audit trails were among the barriers identified.

Walker argued that the quality of the original record becomes increasingly important as AI systems consume and reuse information across multiple business processes.

Walker’s firm’s analysis of banking workflows shows transcripts may be used for customer-feedback analysis, employee training, complaint reviews, investigations and strategic planning.

If transcription errors are repeated or left uncorrected, they risk distorting broader patterns, making a recognition error appear to be a genuine customer trend.

From a quality assurance perspective, this shifts the focus from testing individual AI components to testing complete AI-enabled workflows.

Regression testing, traceability and validation of AI-generated outputs become increasingly important whenever banks update speech-recognition models, large language models, prompts or knowledge bases, ensuring changes do not introduce subtle defects that spread across downstream systems.

Walker said organisations should also consider the level of human review according to the potential impact of each interaction.

“A routine service call, a fraud interview and a legal deposition do not carry the same consequences,” he stresed. “Banks should decide how much human review is required according to how the record will be used.”

As banks accelerate the deployment of AI across customer operations, the quality of the underlying data is becoming as important as the capability of the AI itself.

That increasingly means validating not only whether an AI assistant produces an answer quickly, but whether every stage of the workflow continues to deliver reliable, traceable and trustworthy outcomes.


NEXT MONTH

REGISTER TODAY – SIMPLY CLICK HERE


Why not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

SIGN UP HERE TODAY


REGULATION & COMPLIANCE

Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.


READ MORE


QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS