As banks accelerate their adoption of AI-driven development and increasingly complex digital platforms, the pressure on QA and software testing teams to balance speed, resilience and regulatory compliance has never been higher.
Prince Kohli, President and CEO of Sauce Labs, has a front-row seat to this shift, working with financial institutions grappling with everything from generative AI risk to audit-ready testing frameworks.
In this exclusive interview with QA Financial, Kohli outlines where AI is delivering real gains in quality engineering, where it introduces new vulnerabilities, and why human oversight remains central in regulated environments.
Kohli begins by pointing to three areas where AI is already reshaping quality engineering: “AI genuinely accelerates three critical long-poles of quality engineering right now: test authoring, failure analysis, and coverage at scale.”
He explains that AI can now translate high-level product requirements into executable tests, noting that “AI can take a product manager’s high-level specifications from Jira or a design from Figma and generate executable test cases automatically.”
The impact is particularly evident in failure analysis. With Sauce Labs running “over eight billion tests” on its platform, Kohli says the challenge has shifted from generating data to interpreting it.
“The problem is that interpreting data has become specialised knowledge,” he says, adding that AI now “surfac[es] problem areas and root causes in plain language rather than requiring someone to spend hours combing through multitudes of logs.”
However, Kohli is clear that the risks are amplified in financial services. “When you remove human judgment from the loop entirely, you create a governance gap that your regulatory framework will eventually find,” he warned.

“The problem is that interpreting data has become specialised knowledge.”
– Prince Kohli
In banking, Kohli argues, that gap can have far-reaching consequences: “A transaction flow with a compliance dependency, a behavior that needs to be reconstructable for audit, AI does not inherently know those things matter.”
The cost is not just technical failure but regulatory exposure, from “a customer-facing failure on a critical payment” to “an audit that cannot reconstruct what was actually tested before release,” he added.
For Kohli, the distinction between success and failure lies in how institutions position AI. “The organisations that get this right treat AI as an accelerant for human experts. The ones that struggle automate away the oversight.”
Regulation
The tension between speed and control is also central to how testing platforms must evolve under regulatory frameworks such as DORA. Kohli pushes back on the idea that compliance and velocity are inherently at odds.
“Compliance and speed only appear to be in opposition when your testing infrastructure is the bottleneck,” he says, arguing that “modernise the testing layer and you remove that bottleneck without sacrificing the audit trail.”
What banks increasingly need, he continues, is “verifiable intent.” Rather than simply proving tests passed, platforms must demonstrate purpose, execution and traceability.
“Not just ‘we ran tests and they passed’ but ‘here is the intended behavior we were validating, here is the test that verified it, and here is the record that it ran’,” he says.

This level of traceability, Kohli adds, “is what satisfies DORA and operational resilience requirements, and it needs to be built into the platform, not retrofitted.”
Kohli points to Sauce Labs’ own compliance credentials, including FSQS certification and standards such as SOC 2 Type II and ISO 27001, as part of enabling this visibility.
The goal is to give risk and compliance teams insight into “what was validated, when, and with what confidence level, not just whether tests passed.”
In regulated industries, where data retention can stretch for years, he says “intent-driven testing with associated audit trails makes it much easier to draw the line from purpose to execution and results.”
The challenge becomes more acute as financial platforms grow more complex, spanning mobile, APIs, microservices and third-party integrations.
Kohli describes a familiar frustration among QA teams: “Whatever you wrote yesterday breaks tomorrow, for no reason that feels fair.”
Changes in browsers, dependencies or upstream services can cause widespread test failures, often unrelated to the application itself, he adds.
The root cause, he argues, is that tests are too tightly coupled to implementation. “They were describing how it did it at one specific moment in time,” he says.
“Regulatory accountability sits with people, not systems.”
– Prince Kohli
In fast-moving environments, this brittleness leaves QA teams “permanently behind,” as Kohli puts it.
The shift, he continues, is toward intent-driven testing. “When a test is anchored to intended behavior rather than a specific code path, it does not break every time the underlying plumbing changes. It becomes self-healing by design.”
Combined with cloud infrastructure capable of running large-scale parallel tests across real-world environments, this approach allows institutions to close what Kohli calls the “gap between release velocity and verification velocity.”
GenAI on the rise
Generative AI introduces a further layer of complexity as traditional testing assumptions no longer hold, Kohli explains.
“Traditional testing was built on a simple contract: give the system an input, get a predictable output… Generative AI components do not,” he elaborates.
The variability of outputs is inherent, not a defect, but it undermines deterministic testing models. For financial institutions, the key risk lies in ensuring models behave within acceptable bounds across real-world conditions.
“A model that behaves within acceptable bounds in testing needs to be verified to stay within those bounds across the full distribution of real-world inputs, including adversarial ones,” Kohli says.
This includes protecting against attempts to “extract sensitive data or manipulate a financial output,” he stresses.
Kohli also highlights the need for continuous validation. “The guardrails you validate against one model version need to be re-verified every time that model updates,” Kohli notes, stressing that this is not a rare occurrence in production.
The organisations that succeed, he argues, treat AI systems like any other high-risk dependency, with “continuous behavioral validation, not point-in-time sign-off.”
“Human judgment does not disappear, what changes is where it is applied.”
– Prince Kohli
Despite growing hype around “autonomous testing,” Kohli is careful to define its limits.
“The word ‘autonomous’ sometimes gets used carelessly and it creates the wrong expectations,” he says. While automation can take over discovery, execution and insight generation, “human judgment does not disappear, what changes is where it is applied.”
AI removes the manual burden of writing and maintaining tests and analysing outputs, but the critical decisions remain human-led.
“What remains is the work that actually requires human judgment: defining the intent, evaluating the risk, making the release decision,” he says, describing a model where engineers become “a smart supervisor of agents.”
In banking, that oversight is non-negotiable. “Regulatory accountability sits with people, not systems,” Kohli adds.
Role of data
Finally, Kohli turns to one of the most persistent challenges in QA: data overload. “For many years now, the primary problem has no longer been data,” he says. Instead, organisations are “drowning in signals they do not know how to interpret.”
Different stakeholders, such as engineers, CTOs, compliance officers, all require different insights from the same test runs, yet are often forced into what Kohli describes as “their own tough and complex archaeology.”
AI, he argues, changes that dynamic. A well-trained agent can interpret results “in their massive context,” surfacing root causes, tailoring insights to different roles, and identifying emerging risks.
“A single test failure is an event,” he says. “A pattern of failures… that is a risk signal.” The data to detect these patterns already exists, but the value lies in surfacing them quickly enough to act.
Kohli closes by situating this shift within a broader transformation in software engineering. “The traditional methods… are not being incrementally improved. They are being replaced,” he shares, describing a future where quality is “autonomous, instantaneous, closed loop and stable.”
For financial institutions, the implications are stark. “The question is not whether this change is coming,” he concudes.
“It is whether you are the team that gets ahead of it or the one that is still explaining to your board in two years why your quality and release velocity are both half what your competitors are achieving.”
QA FINANCIAL EVENTS


Why not become a QA Financial subscriber?
It’s entirely FREE
* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *
REGULATION & COMPLIANCE
Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.
READ MORE
- Inside Rabobank: Engineering resilience by design
- Can AI agents finally automate data testing?
- Continuous testing drives DORA compliance
- Why software testing may face a major rethink
- Buy or build? AI rewrites software testing for banks
WATCH NOW

QA FINANCIAL PODCASTS



