As hospitals, insurers and digital health providers race to embed generative AI into patient services and claims systems, a new risk is emerging that goes far beyond software bugs or data breaches: the silent spread of medical misinformation.
For QA and testing teams in healthcare, where system accuracy can be a matter of clinical safety, the challenge is no longer just about functional performance, it’s about accuracy itself.
AI chatbots are increasingly being used to answer patient questions, summarise medical histories and support administrative workflows, a topic that will also feature prominently at the upcoming QA Financial Healthcare & Insurance Forum London 2025, next month in the British capital.
Yet, despite their fluency and speed, these systems can generate factually wrong or misleading statements with unsettling confidence.
This so-called ‘hallucination’ effect has become one of the biggest testing frontiers in the health-tech sector, demanding new QA frameworks that can detect and mitigate false clinical outputs before they reach doctors or patients.
Now, new research from the Icahn School of Medicine at Mount Sinai, one of the United States’ foremost academic medical centres and a global leader in AI-driven health innovation, has put data behind those concerns.
Mount Sinai, home to the Windreich Department of Artificial Intelligence and Human Health, has published one of the first large-scale investigations into how easily large language model (LLM) chatbots can ‘run with’ fabricated medical details.
The findings, published recently in Nature Digital Medicine, provide a stark warning for healthcare QA specialists: even a single piece of false input can cascade into a confident but completely fictional answer.
“Even a single made-up term could trigger a detailed, decisive response based entirely on fiction.”
– Eyal Klang
The study team, led by Eyal Klang, MD, Chief of Generative AI in the Windreich Department, tested six leading language models using 300 clinical vignettes, each containing one invented term, such as a non-existent lab test or disease.
The results revealed hallucination rates between 50 and 82 per cent, with chatbots often elaborating on the fake details as if they were genuine.
“Our goal was to see whether a chatbot would run with false information if it was slipped into a medical question, and the answer is yes,” Dr Klang explained. “Even a single made-up term could trigger a detailed, decisive response based entirely on fiction.”
When the researchers introduced a short ‘safety reminder’ prompt, a simple instruction for the model to verify information, the rate of misinformation fell to around 44 per cent. The improvement, while meaningful, showed that mitigation requires deliberate design rather than parameter tweaking.
“We also found that the simple, well-timed safety reminder built into the prompt made an important difference, cutting those errors nearly in half,” Dr Klang added.
“That tells us these tools can be made safer, but only if we take prompt design and built-in safeguards seriously,” he stressed.
Implications for QA
For QA and testing teams, the implications are clear. Traditional regression or integration testing will not catch these forms of AI error, because they are semantic rather than syntactic, the outputs look correct but contain falsehoods.
Validating generative systems therefore demands a new layer of quality control: adversarial testing that deliberately introduces flawed or misleading inputs to observe whether the system recognises uncertainty or fabricates confidence.
Mount Sinai’s findings also highlight the limits of purely technical mitigation. Adjusting a model’s temperature or deterministic settings did little to reduce hallucinations, suggesting that QA must extend beyond model parameters to encompass prompt engineering, context filters and human oversight.
“The study shows that even well-performing chatbots can generate authoritative but wrong information,” Dr Klang warned, adding that “these tools can be made safer, but only if we take prompt design and built-in safeguards seriously.”
For the healthcare technology community, and the QA professionals who safeguard its systems, the message is unambiguous.
As AI becomes embedded in patient-facing platforms and clinical support tools, safety and accuracy testing must evolve from box-ticking compliance to continuous, adversarial verification.
Mount Sinai’s research makes one thing certain: the future of quality assurance in healthcare will hinge not just on whether AI works, but on whether it tells the truth.

RELEVANT SESSION: AI and Automation for InsurTech: DevOps for connecting payers and providers in an increasingly complex world, with Lee Kivell, global engineering director at Vitality
Wednesday November 26 at 11.30am. More information about the event’s agenda and speakers can be found here.
THIS MONTH

Why not become a QA Financial subscriber?
It’s entirely FREE
* Receive our weekly newsletter every Wednesday * Get priority invitations to our Forum events *

REGULATION & COMPLIANCE
Looking for more news on regulations and compliance requirements driving developments in software quality engineering at financial firms? Visit our dedicated Regulation & Compliance page here.
READ MORE
- Inside Rabobank: Engineering resilience by design
- Can AI agents finally automate data testing?
- Continuous testing drives DORA compliance
- Why software testing may face a major rethink
- Buy or build? AI rewrites software testing for banks
WATCH NOW



