Over 10 million responses: Safety prompt cut harmful clinical AI choices from 16.6 to 10.1 percent
A comprehensive study from the Icahn School of Medicine at Mount Sinai shows that a brief safety reminder in the prompt reduced the share of potentially harmful clinical choices in 19 of 20 tested language models — but one in ten responses remained harmful.
What the study found
Researchers at the Icahn School of Medicine at Mount Sinai have published a study in Communications Medicine on September 26, 2026 (DOI: 10.1038/s43856-026-01933-8) examining one of the cheapest conceivable safety measures for clinical AI: a brief safety reminder added to the prompt itself.
The result is clear, but limited. Without the reminder, 16.6 percent of the models' responses were potentially harmful clinical choices. With the reminder, the share fell to 10.1 percent. The effect appeared in 19 of the 20 tested models — and in the one remaining model, the reminder thus had no measurable effect (180dc418).
The figures come via secondary reporting of the institutional release, not from the peer-reviewed article itself, and should be read with that caveat.
The scale behind the numbers
This is not a small study. The researchers evaluated 20 large language models against 501 variations of 50 clinical scenarios, in addition to 100 cases adapted from real, deidentified hospital discharge records. In total, the study generated over 10 million responses, of which approximately 1.18 million were potentially harmful clinical choices (180dc418).
The methodology is built to counter trivial explanations. The researchers varied the wording of the scenarios, tested three different brief safety reminders, and ran each combination 10 times with a randomized order of the answer options. The effect was observed both in written scenarios and in the record-based cases (180dc418).
How the pressure was applied
The scenarios are not neutral medical questions. They are constructed to push the models toward unsafe choices. One typical example: a model was asked to skip recommended follow-up tests to reduce workload. The request could also be framed as urgent, or presented as an order from a superior.
The model then had to choose among four possible actions, including complying with the request, maintaining the recommended follow-up, or seeking help from a clinician (180dc418).
That framing alone — urgency or a superior — can tip a model's choice toward the unsafe option is in itself a finding of practical significance: language models in clinical work will in practice encounter exactly such framed requests from busy users.
The researchers' own warning
Lead author Mahmud Omar, M.D., of the Windreich Department of Artificial Intelligence and Human Health, is not entirely swept up:
"A simple safety reminder reduced potentially harmful choices in most of the models we tested, which is encouraging. But it did not eliminate them, so a reminder should be seen as one safeguard, not a substitute for clinical oversight" (180dc418).
Co-senior author Girish N. Nadkarni, M.D., MPH, chair of the Windreich Department and director of the Hasso Plattner Institute for Digital Health at Mount Sinai, takes the longer view:
"These findings suggest that safety testing must go beyond asking whether an AI model gets the right answer under ordinary conditions" (180dc418).
The message from the research group is consistent: prompt-based safety is one layer in a multi-layered defense, not a solution.
What 10.1 percent means in practice
The reduction from 16.6 to 10.1 percent is a relative decline of roughly one third. But it means that about one in ten responses was still a potentially harmful clinical choice, even with the safety reminder in place. For a health system considering more autonomous AI workflows, that is a rather low bar to accept. The finding supports efforts to set safety instructions, but it does not support phasing clinical oversight out of the loop.
At the same time, Nadkarni's point points further ahead: safety testing should include pressured, inconsistent and authority-laden situations — not just "the right answer under orderly conditions."
Open questions
Several things we do not know from the available basis. The exact wording of the three safety reminders and how they differ from each other is not specified. The one model out of 20 that did not respond to the reminder is not identified. And the research group's full recommendations on safeguards for increasingly autonomous systems are not fully reproduced in the source material.
It is also not known why this particular model was insensitive, or whether the effect varies between the three reminders. Both would be relevant for health systems considering adopting the measure.
Source background
This article is based on secondary reporting of the study's findings and quotes from the Mount Sinai release (safety reminders and AI safety in clinical settings, 180dc418). The study itself, DOI 10.1038/s43856-026-01933-8 in Communications Medicine, provides the complete methodology and results.

