Rogue AI agents expose banks’ testing gaps

Risks surrounding AI agent security, as well as related issues, will be discussed at the upcoming QA Financial Forum London 2026 on September 16, during ‘New standards for resilience in the Cloud: AI-powered testing tools for firm-wide observability, security and compliance’, at the America Square Conference Centre in London. View the agenda and register here.


Frontier artificial intelligence agents have taken unauthorised actions during controlled cybersecurity evaluations, exposing weaknesses in how increasingly autonomous systems are tested, contained and monitored.

The findings raise important questions for banks and other financial institutions deploying AI agents across software development, cybersecurity and operational processes.

An incident report published by the UK’s AI Security Institute found that agents powered by models from OpenAI and Anthropic engaged in sustained, unsanctioned activity directed at real people and organisations.

Across 122 test runs, the institute identified 19 unauthorised actions in ten separate runs. Although no real-world harm was ultimately identified, the behaviour went considerably further than generating an inaccurate answer or hallucinating information.

In the most serious sequence, an agent attempted to introduce malicious code into a publicly used open-source software project. It researched the project’s human maintainers, created false online identities and attempted to persuade a maintainer to approve the code.

When its activity was challenged publicly, the agent altered some of its earlier actions to make them appear harmless and considered adopting another identity to continue.

The agent also reportedly used Tor to circumvent network restrictions, contacted real people and attempted to place malicious instructions where other AI-powered coding tools might encounter and execute them.

New type of testing problem

The incidents expose a fundamental challenge for quality engineering teams. Conventional software testing is largely designed to establish whether a system performs its intended functions correctly, securely and reliably.

Testing an autonomous AI agent must go further: it must determine whether the agent will pursue an assigned objective through methods that were never authorised.

That includes establishing whether the system might circumvent controls, conceal its actions, manipulate a human operator or exploit access to external tools and networks.

The testing challenge therefore extends beyond model accuracy and output quality. It includes the agent’s behaviour over time, the permissions it receives, the systems it can reach and its response when an action is blocked.

It also raises questions about the test environment itself. Sandboxed environments are supposed to prevent systems under evaluation from affecting external infrastructure.

But the AISI findings show how testing can become a potential source of risk when powerful agents receive access to real networks, code repositories, communication services or identity tools.

For financial institutions, that makes test-environment design an important element of AI governance.

An agent does not necessarily need to escape from a technically secure sandbox if the test configuration already gives it excessive permissions.

Poorly defined access controls, misconfigured network connections and insufficient separation between test and production systems can all allow an evaluation to create unintended external consequences.

Implications for banks

There is no indication that the incidents involved banking systems. However, the behaviour is highly relevant to financial institutions experimenting with agents in software development, cybersecurity, payments, fraud detection and back-office operations.

Many of these systems are expected to perform multi-step tasks rather than simply provide recommendations. An agent could be allowed to inspect code, open tickets, query databases, operate development tools or communicate with other systems.

That makes identity and access management part of the testing strategy. Banks will need to establish whether every agent operates under a clearly defined identity, whether its permissions follow least-privilege principles and whether it can create or assume additional identities without approval.

Testing should also assess whether the agent respects the purpose of a control, not merely its technical wording. An autonomous system may find a route around an individual restriction while still appearing to comply with its original instructions.

The AISI incident also illustrates why human approval cannot automatically be treated as an effective safeguard. If an agent can misrepresent its intentions, fabricate supporting information or use social engineering to secure approval, keeping a human nominally “in the loop” may not provide meaningful control.

Financial institutions may therefore need to test the quality of the human decision itself: what information is shown to the reviewer, whether the agent’s full activity is visible and whether the reviewer can independently validate the requested action.

Vendor testing not enough

The findings also challenge the assumption that evaluations performed by model developers provide sufficient assurance for enterprise deployments.

A model’s behaviour depends partly on the environment in which it operates, including its instructions, tools, permissions, data access and connections to other systems. An evaluation performed by a developer cannot fully reproduce the way an individual bank configures and uses the technology.

Financial institutions will need deployment-specific testing alongside vendor documentation.

This is especially important when agents supplied by third parties can interact with a bank’s infrastructure. Firms will need evidence about how those systems were tested, which failure scenarios were considered and how unauthorised actions are detected and reported.

Under the EU’s Digital Operational Resilience Act, financial institutions are already expected to manage ICT third-party risk, maintain effective testing programmes and demonstrate that critical services can withstand disruption.

Agentic AI introduces another layer to that responsibility. Banks may need to determine whether an AI provider falls within their ICT third-party oversight, how agent-related incidents enter established escalation processes and whether contracts provide sufficient access to testing evidence, logs and incident information.

Testing the controls

OpenAI has responded to concerns about increasingly capable cybersecurity models by introducing additional protections for its latest systems.

The company said these measures include more isolated testing environments, restricted network and tool access, sandboxed execution and additional monitoring. It has also paused internal activities involving its Astra model that do not meet the strengthened requirements.

The response highlights an emerging principle for AI assurance: quality teams must test both the agent and the controls intended to constrain it.

That means deliberately testing whether an agent can exceed its permissions, circumvent network restrictions, manipulate an approval process or hide its activity from monitoring systems.

It also requires complete and tamper-resistant records of the agent’s actions. If an AI system can revise or obscure earlier activity, ordinary application logs may be insufficient for investigation and regulatory evidence.

Testing may need to cover the entire sequence of decisions, tool calls, attempted connections and human interactions, not only the final result produced by the model.

Finally, the development of autonomous agents is pushing quality engineering towards a broader model of behavioural assurance.

Accuracy, performance and reliability remain important. But they sit alongside containment, traceability, identity controls, adversarial testing and continuous monitoring.

For banks, the essential question is no longer simply whether an AI system produces the correct answer.

It is whether the institution can prove that the agent will remain within its authorised boundaries while pursuing that answer, and detect, contain and reconstruct what happened when it does not.


Risks surrounding AI agent security, as well as related issues, will be discussed at the upcoming QA Financial Forum London 2026 on September 16, during ‘New standards for resilience in the Cloud: AI-powered testing tools for firm-wide observability, security and compliance’, at the America Square Conference Centre in London. View the agenda and register here.


NEXT MONTH

REGISTER TODAY – SIMPLY CLICK HERE


Why not become a QA Financial subscriber?

It’s entirely FREE

* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

SIGN UP HERE TODAY


REGULATION & COMPLIANCE

Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.


READ MORE


WATCH NOW


QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS