As artificial intelligence moves deeper into financial operations, quality engineering teams are facing a new challenge: how to test systems that do not simply execute rules, but reason, interpret context and make decisions autonomously.
For banks, insurers and other regulated financial institutions, the stakes are particularly high. AI-driven automation promises major efficiency gains across functions such as accounts payable, expense auditing and transaction validation, but it also introduces new risks around explainability, governance and systemic failure.
In an interview with QA Financial, AppZen founder and CEO Anant Kale explained how the company approaches quality engineering for agentic AI systems used in finance workflows.
Drawing on his extensive background in software engineering, Kale argued that AI models powering financial decision-making must be treated with the same discipline as mission-critical software systems, with rigorous testing, continuous monitoring and strong governance frameworks.
Kale pointed out the foundations of AppZen’s approach were established well before the recent surge in generative AI and agentic AI technologies.
“From the very beginning, we focused on automating finance processes that require human reasoning, domain knowledge, and interpretation of unstructured data like documents, language, and policies,” he stated.
That shift from rule-based automation to AI-driven decision systems significantly changes the role of testing and quality assurance.
“My software engineering background shaped one core belief: finance automation cannot tolerate uncertainty,” Kale elaborated.
“This is not like generating marketing copy or emails where a slightly imperfect result is acceptable. In finance, you cannot approve the wrong amount, apply incorrect accounting, or miss a compliance issue.”

“Rule-based automation can only go so far. Once decisioning is involved, you need AI models.”
– Anant Kale
As a result, AppZen built its AI platform with a strong emphasis on model validation and reliability. “That is why we treated AI models the same way you would treat mission-critical software systems.”
“From day one, we rigorously tested model outcomes using precision and recall, validated them against real transaction sets, and compared model decisions with human judgment.”
Kale saidt that “over time, this evolved into a formal AI reliability and trust program that governs how models are trained, tested, deployed, and monitored in production.”
“Our responsibility is not just building models, but delivering outcomes that meet or exceed human-level accuracy, consistently and transparently.”
Testing decision-making systems
For QA teams accustomed to testing deterministic systems, Kale said the emergence of agentic AI introduces a fundamentally different testing paradigm.
“The biggest change is that you are no longer testing deterministic logic,” he explained. “You are testing decision-making systems that reason, interpret context, and act autonomously.”
“With rule-based systems, you validate whether a rule fired correctly. With agentic AI, you must validate whether the reasoning itself is correct, explainable, and aligned with policy.”
That means QA shifts from checking individual rules to validating decision quality across thousands or millions of real-world scenarios, Kale continued.
He stressed this requires testing approaches that cover both scale and precision. “This requires testing both breadth and depth,” he said.
“Breadth to understand how many use cases a model can handle across industries, countries, and languages. Depth to ensure decisions meet strict accuracy thresholds.”
“This testing is continuous, not one-time, because models learn and adapt based on feedback and outcomes.”
Governance realities
Financial institutions are under growing pressure to improve operational efficiency while maintaining strict controls. Kale believes AI-driven automation is already delivering measurable improvements in certain areas.
“The most immediate gains are in high-volume, judgment-heavy processes like accounts payable, expense auditing, and transaction validation,” he shared.
“These are areas where humans traditionally rely on sampling and post-audit controls, which inherently miss errors.”
“Agentic AI enables 100 percent review of every transaction, with consistent application of policy and controls. When humans and AI work together, error rates are significantly lower than human-only processes,” he noted.
However, Kale warned that expectations around AI autonomy often outpace practical realities.
“Where expectations still outpace reality is in assuming AI can operate without governance,” he argued. “No system, human or AI, is 100 percent accurate.”
“The value comes from combining AI scale with human oversight, strong controls, and continuous monitoring. Organizations that skip that foundation risk disappointment.”

As AI becomes increasingly embedded in financial decision-making, Kale pointed out that explainability and auditability must become core components of quality engineering strategies.
“Trust requires explainability, auditability, and regulatory defensibility,” he said. “Every AI-driven decision includes a decision trace that shows how the outcome was reached.”
“This includes what SOP was followed, what information was extracted from documents, what reasoning steps were applied, and why a specific decision was made.”
Kale went on to add that “these SOPs are created and approved by finance teams to reflect how they would reason manually, and the AI agents execute them consistently.”
This approach can actually improve audit transparency compared with traditional human processed.
“In practice, this means AI decisions can be more auditable than human decisions,” he explained. “With humans, you often only know that someone approved something. With AI, you can see the full reasoning path, which is critical for compliance, audits, and regulators.”
Managing systemic AI risk
As financial institutions automate more operational processes, the risk landscape also shifts.
“As automation increases, the risk shifts from isolated human mistakes to potential systemic issues,” Kale said. “That makes continuous monitoring essential.”
He firmly believes “testing cannot stop at deployment. You need ongoing measurement of live decisions, sampling across workflows, industries, and geographies, and human validation of outcomes.”
He also emphasised the importance of escalation mechanisms that prevent AI systems from making uncertain decisions autonomously.
“Equally important is escalation,” Kale said. “AI systems must be designed to recognie uncertainty and request human input rather than forcing a decision.”
“That combination of continuous monitoring and human-in-the-loop escalation is how you prevent systemic failure.”
“Testing cannot stop at deployment. You need ongoing measurement of live decisions.”
– Anant Kale
Regulatory scrutiny of AI governance is also intensifying globally, particularly in the financial services sector.
“Regulation will reinforce what responsible teams already know: AI systems must be governed, explainable, and resilient,” Kale said.
He pointed out that regulators will increasingly expect clear documentation of how decisions are made, how models are tested, how performance is monitored over time, and how organizations respond to drift or anomalies.
“This aligns closely with how financial institutions already think about model risk management and operational resilience,” Kale explained.
“Agentic AI platforms that embed governance, auditability, and monitoring at the core will be well positioned. Systems that treat AI as a black box will struggle to meet regulatory expectations.”
QA teams as strategic enablers
Looking ahead, Kale believes quality engineering teams will play a much more strategic role as AI becomes embedded across financial operations.
“QA teams need to evolve from point-in-time testers to continuous governors of AI systems,” he said.
“Not all AI use cases require the same level of rigor. Writing emails or generating content is very different from approving financial transactions.”
Kale elaborated that “as AI moves into core operational decision-making, QA teams become strategic enablers by defining accuracy thresholds, monitoring live performance, validating reasoning, and ensuring governance remains intact as systems learn and evolve.”
“This ongoing QA and governance function is what makes AI-driven transformation durable, not just fast.”
Ultimately, Kale believes the future of finance automation will rely on a collaborative model between humans and AI.
“The key point is that agentic AI is not about replacing humans, but about improving decision quality at scale,” he argued.
Kale acknowledged that “humans are not perfect. Traditional finance controls rely on sampling and post-audit processes that inherently miss issues.”
“When AI and humans work together, supported by strong governance, explainability, and continuous quality assurance, the result is higher accuracy, better compliance, and greater trust in outcomes,” he concluded.
“That is the future of finance operations.”
QA FINANCIAL EVENTS


Why not become a QA Financial subscriber?
It’s entirely FREE
* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *
REGULATION & COMPLIANCE
Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.
READ MORE
- Goldman puts AI coding to the test
- How to test AI models that banks do not control
- OpenAI, Filigran and SunTec: the latest vendor and product news
- Sygnum: Testing AI is ‘a measurement problem’
- Banks’ ‘code for all’ push raises testing risks
WATCH NOW

QA FINANCIAL PODCASTS



