Bright
← Living questionsRecord / Health

Google’s AMIE medical AI interviewed real patients before their appointments

A real-patient study of Google’s AMIE found useful pre-visit conversations under physician supervision. Here’s what the results mean for primary care.

Maturity
Emerging, stage 2 of 4
Support
5 sources · paper, institution
Evidence detail
How we know ↓
Conceptual study workflow with three stages: a patient describes symptoms in an AMIE text conversation under separate live physician supervision; the transcript and summary reach the treating clinician; the patient attends an appointment where human clinicians provide care.
explanatory · Original Bright illustration, AI-generated; conceptual study workflow, not actual patients. · Original Bright illustration approved for Bright publication. Not actual patients, an autonomous clinical service, or evidence of improved health outcomes. Source ↗ · View the full-size image ↗

The most useful result in Google’s latest medical-AI study is easy to overlook. In 44 cases where clinicians reviewed the chatbot’s output before seeing a patient, 33 said it helped them prepare for the appointment.

That finding puts a practical possibility into focus: an AI conversation ahead of a visit could help a patient explain what has been happening, then give their clinician a fuller account to work from. In this study, that possibility was tested with people seeking care at a working primary care clinic.

The research involved Google’s Articulate Medical Intelligence Explorer, or AMIE, and Beth Israel Deaconess Medical Center in Boston. Its peer-reviewed publication in The Lancet on October 8 brings a new level of scrutiny to findings first shared publicly in March. The study itself ran from April to November 2025. Google’s publication announcement, BIDMC’s account

What did AMIE do with real patients?

Patients used a text chatbot before a scheduled urgent primary care appointment. AMIE asked about symptoms and medical history, presented possible diagnoses to discuss with their clinician, and produced a transcript and summary for that clinician to review.

The study enrolled 114 adults. Google’s detailed research account says 100 completed an AI conversation; 98 subsequently attended their appointment and formed the analysis group. Each AI conversation had live physician supervision. These were actual care visits, with human clinicians still providing care. Google’s research account, published study abstract

The March preprint adds an important detail about the division of responsibilities. AMIE generated a separate management plan for researchers to evaluate, but that plan was withheld from patients and their providers. The chatbot could discuss possible next steps as topics for the upcoming visit. That distinction matters when interpreting claims about AI treatment planning: the research assessed an output without putting that entire plan into practice. Study preprint

Did the AI help doctors prepare?

The published abstract reports that clinicians returned surveys for 60 of the 98 cases. In 44 of those, they had reviewed the AMIE transcript before the visit. They found it helpful for preparation in 33 cases, or 75%, and said it might have changed their behavior in 25, or about 57%. Those percentages describe the reviewed-transcript cases, rather than the full patient group. Published study abstract

The encouraging idea here is fairly tangible. If a patient’s account is already organized, the appointment can begin with clarifying it and deciding what to investigate. Whether that arrangement reliably improves an appointment is a question for a controlled trial. The present findings capture clinicians’ impressions; they do not establish a measured reduction in workload or improvement in health outcomes.

There is also a workflow problem worth noticing in the survey numbers. In sixteen of the sixty cases with a provider survey, the transcript had not been reviewed beforehand. Producing a useful summary is one task. Getting it to the right person at a time when they can use it is another.

How accurate were AMIE’s possible diagnoses?

Google’s updated research account reports that AMIE’s list contained the eventual diagnosis in 90% of cases. It also reports 75% accuracy within the first three suggestions and 56% for the first suggestion. The reference diagnosis came from chart review eight weeks after the visit. Google’s research account

The March preprint supplies the counts: 88 of 98 within seven possibilities, 73 of 98 within three, and 55 of 98 as the first choice. Its scoring counted the actual diagnosis or a very close suggestion. A reader encountering “90% accurate” deserves those details: a seven-item list offers more opportunities to include the eventual answer than a single prediction. Study preprint

In its research account, Google also describes human evaluations of AI and clinician diagnostic lists and management plans. It reports no statistically significant differences in overall diagnostic-list quality or the safety and appropriateness of management plans. Clinicians scored better on practicality and cost-effectiveness. AMIE had neither access to the medical record nor the ability to examine the patient. These findings concern rated outputs under different information conditions; they cannot establish clinical interchangeability. Google’s research account

What does “no safety stops” mean?

No completed conversation met the study’s predefined criteria for a safety stop. However, supervisors noted one hallucination and added clinical information in five interactions. BIDMC also reports one clinician rating the preparation somewhat harmful, citing concern that mentioning lymphoma could have caused a patient anxiety. “No safety stops” describes a particular monitored endpoint; clinical corrections and concerns still occurred. Published study abstract, BIDMC’s account

The safeguards were substantial. According to the preprint, clinic staff had already screened patients for emergency needs. Eligibility required English, an existing relationship with the clinic and suitable computer access; pregnancy and mental-health chief complaints were excluded. A physician watched the interaction and debriefed the patient afterward. Extending this approach to other populations or less intensive supervision would require fresh evidence. Study preprint

What should the next medical-AI trials measure?

This was a single-center, single-arm feasibility study, without a usual-care control group. Its encouraging conversation ratings and improved attitudes toward AI concern experience and acceptability. BIDMC explicitly says the study was not designed to determine whether AI improves health outcomes. Alphabet funded the research, and co-author Adam Rodman was a visiting Google researcher during part of the study. BIDMC’s account

For a larger comparative trial, the useful questions become more demanding. Do patients leave with a clearer understanding? Are important diagnoses reached sooner? Does the system increase unnecessary testing or anxiety? And does any time saved survive the time spent supervising, correcting and reviewing its output?

The promise of this study lies in bringing those questions into an actual clinic. It offers a plausible starting point for AI assistance: helping a patient’s story reach the clinician before the appointment begins. Establishing the value of that extra conversation is the next piece of work.

How we know5 sources · checked 2026-10-10 · no corrections

Original sources

  1. A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic · Google publication record and final Lancet abstract ↗ · paper
  2. BIDMC research shows safety, quality of AI in primary care · October 9, 2026 ↗ · institution
  3. Exploring conversational diagnostic AI in a real-world clinical study · Google Research, updated October 8, 2026 ↗ · institution
  4. AMIE’s clinical feasibility study published in The Lancet · Google, October 8, 2026 ↗ · institution
  5. Prospective clinical feasibility study · March 2026 preprint, version 3 ↗ · paper

Institutions: Google Research · Beth Israel Deaconess Medical Center

Maturity
Emerging
Source published
2026-10-08
Captured
2026-10-10
Last source review
2026-10-10
Editorial method
AI-assisted source review
Place / relevance
Boston, United States · institution-location

Bright compared this account with the linked original and supporting sources and kept reported, budgeted, projected, and observed claims distinct. Bright did not independently audit the underlying records.

Maturity describes the tested or operational setting. Confidence describes support for the particular claim; one does not determine the other.

Revision & correction history

2026-10-10T12:50:14.970Z · Source-checked in-depth explainer of supervised pre-visit conversations with real primary care patients, preserving study denominators, source-version distinctions and limits.

2026-10-10T12:56:03.361Z · Clarified the standalone Newsroom result sentence using the exact self-contained opening-paragraph wording. Full article body, source citations, image and SEO are unchanged.

No corrections recorded.

Keep exploring

Explore the shared question in another setting. These connections do not imply replication.

An editorial connection recorded in Bright’s evidence catalog.AI is taking on the bottlenecks in breast cancer care ↗Related development · not a replicationAI is taking on the bottlenecks in breast cancer care ↗

Keep looking closer.

See what changed at Bright ↗

Add Bright to your Google Preferred Sources ↗

Suggest a correction · Bright on TikTok