98 patients met Google's AMIE before their doctor visit – no safety incidents, but one hallucination
For the first time, an AI chatbot was tested with real patients in a routine primary care workflow, according to the researchers themselves – and it went ahead without a safety stop. But patients' trust in the chatbot hinged on one thing: that a doctor took over afterwards. The study, published in The Lancet, measured feasibility – not health benefits.
What's new
Researchers at Beth Israel Deaconess Medical Center (BIDMC) in Boston have conducted what they themselves describe as the first prospective real-world study of a patient-facing conversational AI system in primary care. The findings are published in the journal The Lancet (BIDMC/Newswise). The New York Times, which covered the study on October 8, 2026, describes it as "first-of-its-kind" (NYT).
The "first" claim should be handled with precision: it is the researchers' own assertion, not an independently verified ranking. What is remarkable either way is the format – patients used the system from home, ahead of a real urgent-care appointment with their primary care physician, rather than the system being tested only against simulations.
The system in question is Articulate Medical Intelligence Explorer (AMIE), a medical chatbot developed in collaboration between the BIDMC researchers and Google. AMIE remains an experimental research system from Google Research and Google DeepMind and is not in clinical use (Cryptobriefing).
The setup: from the living room to the urgent appointment
The study was conducted between April and November 2025. 114 patients were enrolled, and 98 of them completed both the AI conversation and the subsequent appointment with their primary care physician (BIDMC/Newswise). Secondary sources report the number differently – "nearly 100" at the NYT, "100 real adults" at Cryptobriefing – but the press release's 114/98 is the most precise figure.
Safety oversight was strict. Every conversation was monitored in real time by a board-certified internal medicine specialist, who could intervene if the patient seemed to be at risk, experienced significant emotional distress, asked to end the session, or if the physician identified a potential safety concern – such as the need to clarify symptoms, provide emergency care instructions, or address worry about a possibly serious diagnosis.
The safety results
Across all 98 completed encounters, no conversation required a safety stop. Monitoring physicians identified one hallucination and provided additional clinical clarification in five cases (BIDMC/Newswise).
Those are good numbers for a system speaking directly with patients about symptoms – but they must be read in light of the setup: a physician sat in on every single conversation and could halt anything at any time. The result shows that supervised operation can be conducted safely, not that the system is safe without oversight.
What the doctors got out of it
The primary care physicians were able to read an AI-generated report or summary before the appointment in 44 of the cases. In 75 percent of those cases, physicians said this helped them prepare, and in 57 percent they reported it may have influenced their clinical approach (BIDMC/Newswise). The NYT cites similar figures and quotes resident Peter Brodeur, who compared the chatbot to a very capable medical student or resident (NYT).
There is also a negative case. In a single instance, a clinician reported the interaction was "somewhat harmful," citing concern that a patient may have become anxious after AMIE mentioned lymphoma among possible diagnoses (BIDMC/Newswise). It is a single case – but it points to a real mechanism: a conversational AI that explicitly lists serious diagnoses delivers information to the patient without a clinical frame around it.
Where patients were skeptical
Patients gave the chatbot good marks for its ability to listen, explain information, and create calm. But two areas stood out negatively: trust in the confidentiality of information shared with the system, and trust in the chatbot's honesty and credibility (BIDMC/Newswise).
Adam Rodman, an internist and medical historian at Beth Israel who helped conduct the study, put into words where the trust actually lay: "None of our patients trusted it until they realized that it was handing-off their doctor," he told the NYT (NYT). That may be the study's most important finding: the chatbot was accepted not as an independent actor, but as an extension of a trust relationship that already existed – the one with the primary care physician.
How accurate was AMIE?
There are also figures on diagnostic accuracy, but they do not come from the press release and cannot be verified against the actual Lancet article based on the available source material. Cryptobriefing reports that accuracy was measured after an eight-week review of patient records, which gave physicians time to confirm what patients actually had: the correct diagnosis was in AMIE's top-7 differential in 90 percent of cases, in the top 3 in 75 percent, and as the top-1 candidate in 56 percent (Cryptobriefing). These figures should therefore be read with the caveat that they are so far documented only through secondary coverage.
Key caveats
The researchers themselves stress the most important point: the study measured feasibility, not health outcomes. Nothing in the study can say whether patients got healthier or were treated more correctly because AMIE was involved.
Other limitations: AMIE remains experimental, the study was conducted at a single site with continuous physician oversight, and the patient material of 98 completed encounters is small. And the figures in this coverage come from press materials and secondary reporting – the Lancet article itself was not part of the source material.
What remains
Three questions remain open. First: can the results be replicated at multiple sites, with other patient populations, and without the intensive real-time oversight? Second: whether the study can say anything at all about long-term health outcomes is still unresolved. Third: how should the confidentiality concerns and the trust problem Rodman describes be addressed? If patients accept AI only because a human physician stands behind the handoff, that is a severely limiting condition for scaling.
The study does not prove that AI chatbots belong in primary care. It proves something narrower, but concrete: that the test can be conducted at all – safely, with real patients, in a real workflow. That was the premise that had to be in place before any of the bigger questions could be answered.

