Google’s AMIE medical AI interviewed real patients before their appointments
A real-patient study of Google’s AMIE found useful pre-visit conversations under physician supervision. Here’s what the results mean for primary care.
- Maturity
- Emerging, stage 2 of 4
- Support
- 5 sources · paper, institution
- Evidence detail
- How we know ↓

The most useful result in Google’s latest medical-AI study is easy to overlook. In 44 cases where clinicians reviewed the chatbot’s output before seeing a patient, 33 said it helped them prepare for the appointment.
That finding puts a practical possibility into focus: an AI conversation ahead of a visit could help a patient explain what has been happening, then give their clinician a fuller account to work from. In this study, that possibility was tested with people seeking care at a working primary care clinic.
The research involved Google’s Articulate Medical Intelligence Explorer, or AMIE, and Beth Israel Deaconess Medical Center in Boston. Its peer-reviewed publication in The Lancet on October 8 brings a new level of scrutiny to findings first shared publicly in March. The study itself ran from April to November 2025. Google’s publication announcement, BIDMC’s account
What did AMIE do with real patients?
Patients used a text chatbot before a scheduled urgent primary care appointment. AMIE asked about symptoms and medical history, presented possible diagnoses to discuss with their clinician, and produced a transcript and summary for that clinician to review.
The study enrolled 114 adults. Google’s detailed research account says 100 completed an AI conversation; 98 subsequently attended their appointment and formed the analysis group. Each AI conversation had live physician supervision. These were actual care visits, with human clinicians still providing care. Google’s research account, published study abstract
The March preprint adds an important detail about the division of responsibilities. AMIE generated a separate management plan for researchers to evaluate, but that plan was withheld from patients and their providers. The chatbot could discuss possible next steps as topics for the upcoming visit. That distinction matters when interpreting claims about AI treatment planning: the research assessed an output without putting that entire plan into practice. Study preprint
Did the AI help doctors prepare?
The published abstract reports that clinicians returned surveys for 60 of the 98 cases. In 44 of those, they had reviewed the AMIE transcript before the visit. They found it helpful for preparation in 33 cases, or 75%, and said it might have changed their behavior in 25, or about 57%. Those percentages describe the reviewed-transcript cases, rather than the full patient group. Published study abstract
The encouraging idea here is fairly tangible. If a patient’s account is already organized, the appointment can begin with clarifying it and deciding what to investigate. Whether that arrangement reliably improves an appointment is a question for a controlled trial. The present findings capture clinicians’ impressions; they do not establish a measured reduction in workload or improvement in health outcomes.
There is also a workflow problem worth noticing in the survey numbers. In sixteen of the sixty cases with a provider survey, the transcript had not been reviewed beforehand. Producing a useful summary is one task. Getting it to the right person at a time when they can use it is another.
How accurate were AMIE’s possible diagnoses?
Google’s updated research account reports that AMIE’s list contained the eventual diagnosis in 90% of cases. It also reports 75% accuracy within the first three suggestions and 56% for the first suggestion. The reference diagnosis came from chart review eight weeks after the visit. Google’s research account
The March preprint supplies the counts: 88 of 98 within seven possibilities, 73 of 98 within three, and 55 of 98 as the first choice. Its scoring counted the actual diagnosis or a very close suggestion. A reader encountering “90% accurate” deserves those details: a seven-item list offers more opportunities to include the eventual answer than a single prediction. Study preprint
In its research account, Google also describes human evaluations of AI and clinician diagnostic lists and management plans. It reports no statistically significant differences in overall diagnostic-list quality or the safety and appropriateness of management plans. Clinicians scored better on practicality and cost-effectiveness. AMIE had neither access to the medical record nor the ability to examine the patient. These findings concern rated outputs under different information conditions; they cannot establish clinical interchangeability. Google’s research account
What does “no safety stops” mean?
No completed conversation met the study’s predefined criteria for a safety stop. However, supervisors noted one hallucination and added clinical information in five interactions. BIDMC also reports one clinician rating the preparation somewhat harmful, citing concern that mentioning lymphoma could have caused a patient anxiety. “No safety stops” describes a particular monitored endpoint; clinical corrections and concerns still occurred. Published study abstract, BIDMC’s account
The safeguards were substantial. According to the preprint, clinic staff had already screened patients for emergency needs. Eligibility required English, an existing relationship with the clinic and suitable computer access; pregnancy and mental-health chief complaints were excluded. A physician watched the interaction and debriefed the patient afterward. Extending this approach to other populations or less intensive supervision would require fresh evidence. Study preprint
What should the next medical-AI trials measure?
This was a single-center, single-arm feasibility study, without a usual-care control group. Its encouraging conversation ratings and improved attitudes toward AI concern experience and acceptability. BIDMC explicitly says the study was not designed to determine whether AI improves health outcomes. Alphabet funded the research, and co-author Adam Rodman was a visiting Google researcher during part of the study. BIDMC’s account
For a larger comparative trial, the useful questions become more demanding. Do patients leave with a clearer understanding? Are important diagnoses reached sooner? Does the system increase unnecessary testing or anxiety? And does any time saved survive the time spent supervising, correcting and reviewing its output?
The promise of this study lies in bringing those questions into an actual clinic. It offers a plausible starting point for AI assistance: helping a patient’s story reach the clinician before the appointment begins. Establishing the value of that extra conversation is the next piece of work.
How we know5 sources · checked 2026-10-10 · no corrections
Original sources
- A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic · Google publication record and final Lancet abstract ↗ · paper
- BIDMC research shows safety, quality of AI in primary care · October 9, 2026 ↗ · institution
- Exploring conversational diagnostic AI in a real-world clinical study · Google Research, updated October 8, 2026 ↗ · institution
- AMIE’s clinical feasibility study published in The Lancet · Google, October 8, 2026 ↗ · institution
- Prospective clinical feasibility study · March 2026 preprint, version 3 ↗ · paper
Institutions: Google Research · Beth Israel Deaconess Medical Center
- Maturity
- Emerging
- Source published
- 2026-10-08
- Captured
- 2026-10-10
- Last source review
- 2026-10-10
- Editorial method
- AI-assisted source review
- Place / relevance
- Boston, United States · institution-location
Bright compared this account with the linked original and supporting sources and kept reported, budgeted, projected, and observed claims distinct. Bright did not independently audit the underlying records.
Maturity describes the tested or operational setting. Confidence describes support for the particular claim; one does not determine the other.
Revision & correction history
2026-10-10T12:50:14.970Z · Source-checked in-depth explainer of supervised pre-visit conversations with real primary care patients, preserving study denominators, source-version distinctions and limits.
2026-10-10T12:56:03.361Z · Clarified the standalone Newsroom result sentence using the exact self-contained opening-paragraph wording. Full article body, source citations, image and SEO are unchanged.
No corrections recorded.
Keep exploring
Explore the shared question in another setting. These connections do not imply replication.
