
The International Organization of Securities Commissions (IOSCO) has unveiled a new supervisory toolkit for artificial intelligence in capital markets, placing governance, testing, operational resilience and lifecycle oversight firmly at the centre of how regulators expect banks and financial services firms to deploy AI systems across trading, compliance, surveillance and investment operations.
The toolkit reflects growing concern among regulators that the rapid rollout of Generative AI and emerging Agentic AI systems is accelerating faster than many firms’ governance and quality assurance frameworks can safely absorb.
For QA and software testing teams inside banks, brokerages, exchanges and asset managers, the message is increasingly clear: AI systems can no longer be treated as isolated innovation projects. They are becoming core operational infrastructure requiring continuous testing, validation, monitoring and supervisory evidence.
“The increasing integration of artificial intelligence into capital markets requires supervisors to have practical and proportionate tools to assess emerging risks while supporting innovation and safeguarding market integrity and investor protection,” explained Jean-Paul Servais, who is the board chair of IOSCO, the global standard-setter for the securities and capital markets sector.
It brings together the world’s national securities regulators, such as the SEC in the United States and the FCA in the UK, to establish and promote consistent international standards of regulation and enforcemen.
Operational supervision
The IOSCO toolkit is designed as a practical supervisory framework covering the full lifecycle of AI systems, ranging from traditional machine learning models to Generative AI and emerging agentic architectures.
The report explicitly focuses on how regulators can supervise AI governance, third-party dependencies, disclosure obligations, operational resilience, recordkeeping and model oversight as firms move from experimentation toward scaled deployment.

“The toolkit reflects IOSCO’s continued commitment to strengthening supervisory approaches in response to rapidly evolving technologies,” stressed Hanzo van Beusekom, Chair of IOSCO’S Fintech Task Force and member of the executive board of the Dutch Authority for the Financial Markets.
“It provides practical guidance to help authorities address the risk arising from the growing use of AI systems across financial markets.”
The report comes at a time when many capital markets firms are rapidly integrating AI into trading systems, surveillance operations, client onboarding, fraud detection, investment research, risk management and operational automation.
Madrid-based IOSCO noted that “the use of AI in capital markets presents opportunities and risks that supervisors must understand and address to protect investors, maintain market integrity, and preserve financial stability.”
AI rollouts
The report highlights how many financial institutions have moved beyond limited AI experimentation into broader operational deployment, particularly using Generative AI systems.
According to IOSCO, “the Survey data indicates there is an increase in the speed and scale of AI development and implementation over the last two years, particularly using GenAI.”
That acceleration is now creating direct implications for software testing and QA teams.
The watchdog warns that AI systems introduce risks around “complexity, reduced transparency, third-party dependencies and governance challenges,” while newer GenAI and Agentic AI architectures create additional problems around explainability, reliability, cybersecurity and operational control.
One of the report’s strongest themes is that AI governance cannot rely solely on traditional software controls.
“GenAI systems may be difficult to predict, evaluate, understand, explain, and test,” IOSCO warned.
That creates new pressure on firms to continuously validate AI behaviour after deployment, rather than relying only on pre-release testing.
Supervisory testing enters the AI era
Perhaps the clearest shift in the IOSCO framework is the move toward continuous supervisory engagement around AI systems.

The toolkit is explicitly designed for use during “on-site examinations and inspections,” with regulators expected to review governance structures, testing procedures, validation frameworks, monitoring controls, incident response mechanisms and AI documentation directly.
“The toolkit marks an important milestone, culminating two years of work by the Fintech Task Force,” according to said Lim Tuang Lee, former chair of IOSCO’S Fintech Task Force and assistant managing director, Capital Markets Group, Monetary Authority of Singapore.
“It is a significant step to strengthen supervisory readiness among capital market supervisors and represents IOSCO’s continuous effort to support robust oversight while facilitating responsible innovation,” he argued.
For QA and software testing leaders across financial markets, the direction is unmistakable. AI governance is moving rapidly from a policy discussion into a testing discipline.
Hallucinations and AI failures
A major focus of the toolkit is the risk of hallucinations and unpredictable outputs from GenAI systems.
IOSCO describes hallucinations as “outputs that appear plausible but are factually or logically incorrect or fabricated,” warning that these risks become particularly problematic in financial markets where “trust and credibility are paramount.”
The report outlines several mitigation approaches now attracting regulatory attention, including Retrieval-Augmented Generation, Chain-of-Verification frameworks, multi-agent debate models and stronger human oversight mechanisms.
However, IOSCO also cautions that “these techniques have limitations and are unlikely to fully eliminate the risk of hallucinations.”
That warning is likely to resonate strongly with QA and model validation teams already struggling to test probabilistic AI systems that can produce inconsistent outputs even when given similar prompts.
The toolkit repeatedly emphasises lifecycle testing and ongoing monitoring rather than one-time validation exercises.
IOSCO said firms should maintain “validation and testing policies & procedures and reports, including pre-deployment, post-deployment and ongoing validation, monitoring and testing.”
The report also highlights supervisory concerns around model drift, adversarial attacks, data poisoning, bias, AI system failures, and single points of failure.

Agentic AI
One of the most significant sections of the report focuses on Agentic AI systems, which IOSCO says are becoming increasingly relevant to capital markets oversight.
These systems, typically equipped with planning capabilities, long-term memory and access to external tools or systems, could introduce new operational and cyber risks if not properly governed.
IOSCO warned that “the potential interplay between the components could create the risk of unexpected emergent behaviors or the potential for cascading impacts or failures across interconnected systems.”
The watchdog also highlighted risks involving “collusive behaviors,” “misaligned goals,” and “poorly designed prompts” that could lead to operational or security failures.
For QA teams, that shifts testing requirements far beyond traditional regression testing or deterministic validation.
The toolkit suggests firms may need continuous monitoring, gradual rollouts, resilience and red team testing, rollback plans, and human oversight configurations calibrated to the risk profile of specific AI systems.
Core controls
A central message throughout the IOSCO framework is that governance structures must evolve alongside increasingly autonomous AI systems.
The toolkit repeatedly stresses the need for centralized AI governance, enterprise-wide risk management frameworks, documented policies, AI inventories, independent model validation, board-level oversight and clearly defined human intervention mechanisms.
“Robust governance and risk management frameworks are essential to ensuring responsible and resilient deployment of AI systems into financial products and services,” IOSCO stated.
The report also places heavy emphasis on transparency, explainability and auditability.
According to IOSCO, firms should maintain “AI system decision and usage logs and audit trail documentation, including inputs, outputs, and AI system logic.”
That requirement has direct implications for engineering and QA functions inside banks and capital markets firms, particularly where firms are embedding GenAI into surveillance systems, trading workflows, investment recommendations or operational decision-making.
The report repeatedly warns that supervisors may increasingly expect firms to demonstrate not just that AI systems work, but that they can be governed, explained and interrupted safely when necessary.

IOSCO also linked AI adoption to rising cybersecurity and operational resilience concerns.
The organization warned that “AI-enabled cyber capabilities” could “materially accelerate threat evolution and increase the speed, scope, and scale of existing attack techniques.”
That concern becomes particularly significant in environments where multiple firms rely on the same cloud providers, foundational models or AI vendors.
The report flags concentration risks and outsourcing dependencies as a growing supervisory concern, especially where multiple institutions rely on common AI infrastructure providers.
For firms active in capital markets, this creates a growing requirement for operational resilience testing that extends beyond internal applications into external AI ecosystems and third-party model dependencies.
Supervisory testing enters the AI era
Perhaps the clearest shift in the IOSCO framework is the move toward continuous supervisory engagement around AI systems.
The toolkit is explicitly designed for use during “on-site examinations and inspections,” with regulators expected to review governance structures, testing procedures, validation frameworks, monitoring controls, incident response mechanisms and AI documentation directly.
“The toolkit marks an important milestone, culminating two years of work by the Fintech Task Force,” according to said Lim Tuang Lee, former chair of IOSCO’S Fintech Task Force and assistant managing director, Capital Markets Group, Monetary Authority of Singapore.
“It is a significant step to strengthen supervisory readiness among capital market supervisors and represents IOSCO’s continuous effort to support robust oversight while facilitating responsible innovation,” he said.
For QA and software testing leaders across financial markets, the direction is unmistakable. AI governance is moving rapidly from a policy discussion into a testing discipline.
WHY not become a QA Financial subscriber?
It’s entirely FREE
* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *
READ MORE
- Inside Rabobank: Engineering resilience by design
- Can AI agents finally automate data testing?
- Continuous testing drives DORA compliance
- Why software testing may face a major rethink
- Buy or build? AI rewrites software testing for banks
WATCH NOW

QA FINANCIAL PODCASTS

CLICK HERE TO LISTEN TO OUR EXCLUSIVE CONVERSATIONS



