How do banks test and control AI that acts alone?

Tech Mahindra's Gopal Parasni

Banks will need to move beyond conventional model testing as agentic AI begins making decisions and executing entire financial workflows.

The move from generative AI tools that draft, summarise and assist towards autonomous agents capable of planning work, interacting with multiple systems and completing processes with limited human involvement significantly expands the quality assurance challenge

Testing an underlying model will not prove that an agent can reliably complete a loan application, investigate suspected fraud or rebalance an investment portfolio across a chain of connected banking systems.

A new white paper from Tech Mahindra examined this shift and the controls banks will require. Gopal Parasnis, head of AI and digital transformation for financial services at the company and co-author of the paper, stressed that AI agents “must never operate in complete isolation”, particularly within regulated financial services.

Agentic systems can decide which tools to use, retrieve information from different sources and adjust their actions in response to changing conditions. Banks will therefore need to test not only individual outputs but the complete sequence of decisions, system calls and transactions generated by an agent.

This includes establishing whether agents select the correct data, perform actions in the proper order, remain within their permissions and escalate exceptions when human intervention is required. Testing teams may also need to examine how agents respond to missing, conflicting, manipulated or unexpected information.

The paper identified loan processing, fraud investigation and portfolio rebalancing as examples of workflows that agents could eventually manage from beginning to end.

It also set out potential applications across underwriting, payments, customer onboarding, anti-money laundering, reconciliation and investment management.

Such use cases introduce a substantially higher level of testing risk because a plausible response is no longer sufficient. An agent’s output may initiate a payment, influence a credit decision, change a customer record or trigger a compliance process.

A model could perform accurately in isolation while the overall workflow still fails because of defective orchestration, unreliable APIs, poor-quality data, incorrect permissions or weak integration with legacy banking platforms.

Traceability as a QA requirement

The paper makes auditability a central part of the proposed governance framework. Parasnis said banks should ensure that “every decision made by an agent is traceable and logged for audit”.

This means observability will have to extend beyond conventional system logs. A bank may need to reconstruct which data an agent accessed, which tools it selected, what intermediate decisions it made and why it proceeded, stopped or referred a case to a person.

That evidence will be particularly important when an agent contributes to a regulated decision. The paper uses lending as an example, arguing that when an agent approves or declines a loan, the decision must be explainable and free from bias.

Testing programmes will consequently need to cover fairness, explainability and consistency alongside functional accuracy, security, performance and resilience. They will also need to establish whether the same customer circumstances produce acceptably consistent outcomes across repeated runs and model or prompt updates.

Continuous monitoring is another core part of the framework. Parasnis said responsible AI practices should ensure models are “tested, validated, and monitored” in alignment with established model-risk management practices.

This suggests that pre-production assurance will represent only one stage of the testing lifecycle. Banks will need controls for detecting behavioural drift, declining accuracy, changing data conditions and unexpected interactions between agents after deployment.

Human controls to be tested

The paper advocates a human-in-the-loop model in which people review, approve or override critical decisions. However, the existence of a human approval stage does not by itself demonstrate that the control is effective.

Banks will need to test whether escalation thresholds activate under the right conditions, whether agents can circumvent approval gates and whether reviewers receive enough contextual information to make an informed decision. Failure scenarios should also cover unavailable reviewers, delayed approvals and disagreements between human and automated recommendations.

The paper said banks should clearly define which decisions agents can make independently, when people must become involved and which boundaries agents may never cross.

Those boundaries will need to become executable and repeatable test conditions rather than remaining high-level governance statements. This could require negative testing, adversarial testing and scenario-based validation covering both foreseeable errors and unusual combinations of events.

The researchers recommended that banks begin with lower-risk back-office applications before moving towards customer-facing deployments.

Its proposed implementation roadmap progresses from readiness assessments and controlled pilots to multi-agent workflows, continuous feedback and performance monitoring.

They also promoted its Orion platform and VerifAI framework as tools for developing, governing and validating agentic banking applications.

The paper claimed early agentic AI deployments can produce productivity improvements of between 50% and 60% in selected processes and reduce false positives in fraud detection by 30% to 40%.

However, moving those systems safely from pilots into production will depend on whether banks can test the behaviour of the complete agentic ecosystem rather than simply assess the model at its centre.

As autonomous agents gain access to customer information, banking APIs and transaction systems, software quality becomes part of the control framework itself. Banks will need evidence that agents not only produce credible answers, but take the correct actions, within defined boundaries, under both normal and abnormal conditions.

That makes agentic AI a broader assurance challenge than the current wave of generative AI: banks are no longer merely testing what a system says, but what it decides and does.


NEXT MONTH

REGISTER TODAY – SIMPLY CLICK HERE


Why not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

SIGN UP HERE TODAY


READ MORE


QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS