A second opinion tested across three health systems.
A 249-physician study in Kenya, Indonesia, and the Netherlands tested GPT-4o assistance on clinical vignettes and reported higher guideline-based scores in each setting.
Original sources ↓ · Revision history ↓
Demonstrated · source published 2026-09-09
The human problem
Clinicians must apply changing guidance under time and information pressure, with uneven access to specialist support.
The prior constraint
General-purpose clinical assistants may perform differently across languages, guidelines, and health systems, and many evaluations do not include practicing clinicians across countries.
AI’s actual role
GPT-4o supplied information and reasoning support while physicians retained the task and answer.
The documented result
Among 249 physicians, guideline scores increased by 18 percentage points in Kenya, 10.7 in Indonesia, and 7.2 in the Netherlands. These were vignette scores, not patient outcomes.
Why it may matter
The study asks whether assistance transfers across settings while making the geographic differences visible.
Limitations
The tasks were clinical vignettes, not live care.
Guideline-score gains do not establish safety, diagnostic accuracy, or patient benefit.
One model version and study interface may not generalize to other tools or changing models.
Unresolved questions
Which specialties and case types account for errors or gains?
How does assistance affect time, overreliance, and disagreement in real practice?
Do language and local guideline differences change safety?
Source history & evidence assessment
- Maturity
- Demonstrated
- Claim confidence
- high
- Event date
- 2026-09-09
- Source published
- 2026-09-09
- Captured
- 2026-09-19
- Last source review
- 2026-09-19
- Editorial method
- AI-assisted source review
- Place / relevance
- Kenya, Indonesia, and the Netherlands · global-study
AI-assisted editorial comparison with the cited primary sources, explicit evidence limits, and held alternatives. Publication authorized by the site owner on 2026-09-19; no human source review or independent replication is claimed.
Maturity describes the tested or operational setting. Confidence describes support for the particular claim; one does not determine the other.
Original sources
Impact of LLM assistance on physician decision-making: a multi-country randomized controlled trial ↗ · paper
The Big Unknown: A Journey Into Generative AI's Transformative Effect on Meical Professions ↗ · registry
Institutions: Study sites in Kenya, Indonesia, and the Netherlands
Explore the underlying question
Related developments
Editorial connections between distinct settings and results; these links do not imply replication.
Revision & correction history
2026-09-19 · Initial three-country vignette-study draft.
No corrections recorded.
