Could smaller AI models make software testing easier for banks?

Scott Zoldi

As banks accelerate the deployment of generative AI, one of the biggest gains for software testing teams may come not from increasingly powerful large language models, but from using smaller, purpose-built AI models that are easier to validate, monitor and govern.

At least, that is the view of Scott Zoldi, who works for analytics software provider FICO, whose technology is used by banks, insurers and financial institutions worldwide for areas including credit scoring, fraud detection, decision management and risk analytics.

Zoldi thinks banks should at least test and trial smaller, tailored AI infrastructure as financial firms should move away from relying on general-purpose AI models for business-critical decisions and instead deploy smaller models trained exclusively for specific financial services use cases.

According to the FICO chief analytics officer, that approach not only reduces computing costs but could also simplify software testing, model validation and regulatory compliance in an increasingly demanding supervisory environment.

Zoldi argued that today’s frontier AI models are not yet suitable for many regulated financial applications.

“The open models that are out there today and the global models that are out there today don’t meet the standards of responsible AI,” he explained to website The AI Innovator.

“They need to get to robust AI, explainable AI, ethical AI and auditable AI to ensure that the basic data science practices we’ve been doing for decades are not violated.”

Rather than attempting to validate models trained on vast amounts of internet data, banks could focus on testing models designed around narrowly defined business domains such as lending, fraud detection, anti-money laundering or customer servicing, Zoldi stressed.


“If you’re a financial services firm, you do not need a model that’s built on Spice Girls lyrics.”

– Scott Zoldi

A smaller testing surface makes it easier to perform regression testing, explainability assessments and model validation, while reducing the number of unpredictable outputs that must be evaluated before deployment.

It also simplifies demonstrating evidence of governance to regulators, who are placing increasing emphasis on explainability, auditability and operational resilience in AI systems.

Zoldi believes the scope of the model itself is fundamental to achieving those goals. “If you’re a financial services firm, you do not need a model that’s built on Spice Girls lyrics. You do not need a model to tell you how to change a tire. You need a model that only knows about financial services.”

He also questions the growing enterprise trend of connecting proprietary banking data to large, general-purpose models through retrieval augmented generation (RAG) or extensive fine-tuning.

“A large commoditised model has seen the entire totality of all the data known to mankind,” Zoldi continued. “So how do we think that we’re going to persuade it to ignore everything that it has learned? By prompting correctly? How do you have confidence in that?”

Instead, he advocates building multiple specialised AI models, each trained for a single business task using carefully curated datasets and validated examples rather than historical operational data that may itself contain bias or poor decisions.

Focus on performance

According to Zoldi, these smaller models can outperform much larger general-purpose systems on financial tasks because they are trained exclusively on relevant knowledge.

Perhaps the most interesting proposal from a software testing perspective is FICO’s “trust score” architecture, which effectively creates an additional layer of automated AI quality assurance.

Rather than relying solely on confidence scores generated by the primary AI model, FICO proposes deploying a second AI model that independently evaluates the first before responses are delivered.

Business specialists, compliance professionals, lawyers and risk experts define “knowledge anchors” describing how particular scenarios should be handled. The secondary model then assesses whether responses align with those approved standards before assigning a numerical trust score.

For banks, this effectively introduces continuous AI validation rather than relying solely on testing before deployment.

It also creates a new challenge for quality engineering teams: how should the validation model itself be tested, monitored and regression-tested as the primary AI model evolves?

As organisations increasingly deploy AI to verify other AI systems, software assurance is itself becoming a multi-layered discipline, Zoldi pointed out.

Cost is another important part of FICO’s argument. Smaller, domain-specific models require significantly less computing infrastructure and reduce dependence on expensive third-party AI services.

“That’s not hard. Get your PyTorch,” Zoldi said, adding that “you don’t need a lot of GPUs.” Recalling FICO’s own experience, he said: “It’s not a big investment.”

For financial institutions, however, the savings may extend well beyond infrastructure. Smaller, specialised AI models could also reduce the cost of software testing and model assurance by making validation exercises more targeted, explainability easier to demonstrate and governance evidence simpler to produce.

At a time when regulators around the world are demanding greater confidence in how banks develop, test and monitor AI systems, that combination of lower operating costs and easier assurance could prove just as valuable as any reduction in GPU expenditure.


16 SEPTEMBER IN LONDON

REGISTER TODAY – SIMPLY CLICK HERE


Why not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

SIGN UP HERE TODAY


REGULATION & COMPLIANCE

Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.


READ MORE


WATCH NOW


QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS