As banks and financial services firms accelerate the use of AI-assisted development, autonomous coding tools and large language models inside core systems, software testing and QA teams are being pulled into a new kind of risk conversation.
Traditional functional testing, regression suites and even automated security scanning are increasingly being stretched by systems that behave unpredictably, generate code at speed and operate across sprawling digital estates.
In regulated environments, where outages, data leaks or logic flaws can trigger regulatory scrutiny as well as financial loss, the pressure is mounting on QA and testing functions to move beyond validation and into active risk discovery.
That shift is pushing offensive security testing, specifically human-led penetration testing, closer to the heart of quality and resilience strategies.
It is against this backdrop that Cobalt’s latest research and platform developments are reframing pentesting not as a compliance checkbox, but as a critical extension of modern QA in the age of AI-driven software delivery.
Offensive testing
“The cybersecurity landscape has reached a watershed moment,” Cobalt stated in its latest Responsible AI Imperative white paper, warning that the rapid adoption of AI across the software development lifecycle is outpacing organisations’ ability to manage the risks it introduces.
“AI-powered coding agents are now foundational to the modern development lifecycle, promising unprecedented velocity. However, this rapid transformation has created a critical readiness gap,” the report pointed out, noting that such tools are “often bypassing foundational security principles.”
For QA and testing teams in financial services, the report’s findings point directly to the limits of conventional testing approaches when applied to AI-enabled systems.
Cobalt’s data shows that 32% of findings in LLM pentests are classified as serious, high or critical risk, the highest proportion found across all asset types, while 36% of security leaders admit that the demand for genAI has outpaced their ability to manage its security implications.
Perhaps most concerning for teams responsible for software quality and resilience is what happens after issues are found.
The report highlighted a stark remediation gap, revealing a 21% resolution rate, with “only one-fifth of serious LLM vulnerabilities in Cobalt pentests actually resolved, the lowest rate among all penetration test types.”
Why automation alone is not enough
The white paper drew a clear line between automated testing tools and the realities of modern attack surfaces, particularly where AI is involved.
“Automated scanners cannot catch the non-deterministic logic flaws or the creative chaining of exploits unique to LLMs,” it stated, arguing that this class of risk demands a more exploratory, human-driven testing mindset.
For QA teams already grappling with complex integrations, legacy platforms and regulatory deadlines, the message is blunt.
To address these blind spots, organisations must “integrate offensive security testing into every stage of the AI lifecycle,” and explicitly “mandate human-led pentesting.”
That emphasis aligns closely with broader shifts in quality engineering, where testers are increasingly expected to challenge assumptions, simulate failure scenarios and stress systems in ways that mirror real-world behaviour rather than idealised test cases.
Pentesting fit for modern workflows
Alongside its research, Cobalt has been evolving its offensive security platform to make pentesting easier to initiate and more tightly aligned with development and testing cycles.
Jason Lamar, Senior Vice President of Product at Cobalt, said recently that platform updates were designed to remove friction from how organisations engage with pentesting.

“These innovations mark the next chapter in the evolution of offensive security services,” Lamar explained, describing changes aimed at improving accessibility, visibility and actionability for teams under pressure to move faster without increasing risk.
Lamar framed the goal in deliberately practical terms, saying Cobalt wants to make pentesting “as simple as ordering a pizza,” reflecting a push to embed offensive testing more naturally into day-to-day software delivery rather than treating it as a specialist, last-minute exercise.
“So we are building toward a future where pentesting is continuous, deeply integrated into development workflows, and backed by data that drives real security outcomes, not just compliance,” he said, pointing to a broader redefinition of pentesting as part of ongoing quality assurance rather than a standalone security event.
For banks and financial services firms navigating AI adoption under intense regulatory scrutiny, the convergence of QA, resilience and offensive security testing is becoming harder to ignore.
The firm’s research and platform direction both underscore a common conclusion: in an era of autonomous systems and accelerating change, finding flaws before attackers do may depend less on automation alone, and more on bringing human-led pentesting firmly into the QA toolkit.
COMING IN 2026


Why not become a QA Financial subscriber?
It’s entirely FREE
* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *
REGULATION & COMPLIANCE
Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.
READ MORE
- Inside Rabobank: Engineering resilience by design
- Can AI agents finally automate data testing?
- Continuous testing drives DORA compliance
- Why software testing may face a major rethink
- Buy or build? AI rewrites software testing for banks
WATCH NOW

QA FINANCIAL PODCASTS



