Bright
← Living questionsRECORD / Health

A second opinion tested across three health systems.

A 249-physician study in Kenya, Indonesia, and the Netherlands tested GPT-4o assistance on clinical vignettes and reported higher guideline-based scores in each setting.

Original sources ↓ · Revision history ↓

Demonstrated · source published 2026-09-09

The human problem

Clinicians must apply changing guidance under time and information pressure, with uneven access to specialist support.

The prior constraint

General-purpose clinical assistants may perform differently across languages, guidelines, and health systems, and many evaluations do not include practicing clinicians across countries.

AI’s actual role

GPT-4o supplied information and reasoning support while physicians retained the task and answer.

The documented result

Among 249 physicians, guideline scores increased by 18 percentage points in Kenya, 10.7 in Indonesia, and 7.2 in the Netherlands. These were vignette scores, not patient outcomes.

Why it may matter

The study asks whether assistance transfers across settings while making the geographic differences visible.

Limitations

The tasks were clinical vignettes, not live care.

Guideline-score gains do not establish safety, diagnostic accuracy, or patient benefit.

One model version and study interface may not generalize to other tools or changing models.

Unresolved questions

Which specialties and case types account for errors or gains?

How does assistance affect time, overreliance, and disagreement in real practice?

Do language and local guideline differences change safety?

Source history & evidence assessment
Maturity
Demonstrated
Claim confidence
high
Event date
2026-09-09
Source published
2026-09-09
Captured
2026-09-19
Last source review
2026-09-19
Editorial method
AI-assisted source review
Place / relevance
Kenya, Indonesia, and the Netherlands · global-study

AI-assisted editorial comparison with the cited primary sources, explicit evidence limits, and held alternatives. Publication authorized by the site owner on 2026-09-19; no human source review or independent replication is claimed.

Maturity describes the tested or operational setting. Confidence describes support for the particular claim; one does not determine the other.

Original sources

Impact of LLM assistance on physician decision-making: a multi-country randomized controlled trial · paper

The Big Unknown: A Journey Into Generative AI's Transformative Effect on Meical Professions · registry

Institutions: Study sites in Kenya, Indonesia, and the Netherlands

Explore the underlying question

Related developments

Editorial connections between distinct settings and results; these links do not imply replication.

Faster to start, but not the coach everyone wanted.

Revision & correction history

2026-09-19 · Initial three-country vignette-study draft.

No corrections recorded.