Bright
← Living questionsRecord / Health

AI finds heart risk clues in overnight sleep tests

A new study suggests sleep-test recordings could help flag future heart problems. The next challenge is proving that the extra warning improves care.

Maturity
Emerging, stage 2 of 4
Support
7 sources · institution, paper
Evidence detail
How we know ↓
An empty sleep-study room with a bed, cabinets and equipment on a bedside table.
contextual · Illustrative image: a sleep-study room at Naval Medical Center San Diego, photographed on March 25, 2021. The research discussed here used records from other hospitals. U.S. Navy photo by Mass Communication Specialist 3rd Class Jake Greenberg, via DVIDS (public domain). The appearance of U.S. Department of War (DoW) visual information does not imply or constitute DoW endorsement. · Public domain · U.S. Navy photograph; DVIDS public-use conditions and non-endorsement notice retained. Not this research team, its hospital sites or participants, the AI system, or demonstrated improved patient outcomes. Source ↗ · View the full-size image ↗

An overnight sleep test gathers hours of information about someone’s breathing, brain activity and heartbeat. Doctors use it to investigate problems such as sleep apnea, when breathing repeatedly stops and starts. That night of monitoring may also contain clues to a patient’s longer-term health. NHLBI’s guide to sleep studies

Research highlighted by the National Institutes of Health on October 9 explores one possibility: using artificial intelligence to extract information about future heart risk from the ECG, a recording of the heart’s electrical activity, already collected during a sleep study. The system combines that signal with information about sleep stages, the phases a person moves through while sleeping. NIH research announcement

The attraction is practical. If the approach eventually proves useful, patients undergoing these tests could get more value from the same night of monitoring. For now, the evidence comes from past records. The study does not establish that using its predictions improves patients’ health. The study in SLEEP

What the researchers found

The model was fine-tuned using 15,809 patients at Massachusetts General Hospital. Researchers then tested it on separate groups: 9,810 patients at Emory University Hospital and 12,576 at Beth Israel Deaconess Medical Center. It used a single ECG channel and expert-labeled sleep stages. NIH research announcement

The study targeted outcomes over 10 years. Adding the model’s output to clinical and sleep-related risk factors improved discrimination for atrial fibrillation, heart failure and death from any cause. Discrimination means distinguishing people who later experience an outcome from those who do not. The model did not improve that distinction for stroke or heart attack. The study in SLEEP

Atrial fibrillation is an irregular heart rhythm. Heart failure means the heart cannot pump enough blood to meet the body’s needs. Both are serious conditions, but a prediction of greater future risk does not tell a person that they have either condition now. NHLBI on atrial fibrillation, NHLBI on heart failure

Why testing at other hospitals matters

The separate hospitals are an important part of the result. A model can appear impressive when it learns patterns peculiar to the place that supplied its training data. Testing on patients from other hospitals asks a tougher question: does the useful signal survive a change of setting?

Here, the encouraging interpretation is that the method deserves investigation beyond its original development site. It still needs testing across the range of people, equipment and working conditions where clinicians might use it. General principles for medical AI explicitly call for representative patients and independent test data. Medical AI development principles

The study population consisted of hospital sleep-study patients. It does not establish performance for everyone who wears a watch to bed or uses a home sleep device. The study in SLEEP

That boundary matters to the promise. A useful tool for a defined group can still be valuable. Reaching people outside that group is a separate research question.

A warning needs to mean something

Imagine a future clinic receiving a risk score alongside a sleep report. Before acting, staff would need to know how accurately the number describes patients like theirs. If a system assigns a 10 percent risk to a group, roughly one in ten should experience the specified outcome within the stated period. This is calibration, a different question from whether the system ranks people in the right order.

The paper reports calibration over a six-year follow-up period and retrospective decision-curve analyses. Those checks do not establish the effects of using the model in care. The study in SLEEP

For this application, a useful next study would specify an actual decision: which scores trigger a review, who conducts it and what information they consider. It should measure how many people receive worthwhile follow-up and how many are sent through unnecessary appointments. A mathematically better prediction could still be difficult to use if the next step is vague.

The limits of looking backward

The researchers identified cardiovascular diagnoses through medical-record codes, without direct clinical confirmation. These records can contain errors. They also acknowledged the difficulty of separating future atrial fibrillation from previously unrecognized episodes. The study in SLEEP

That raises a practical question for further testing: what, exactly, is the score picking up in each patient? A signal of existing illness and a warning of future illness can lead to different conversations and different investigations. Clinicians need enough information to interpret the result responsibly.

Fairness deserves similarly specific scrutiny. Future evaluations should ask whether errors concentrate in particular age groups, sexes or racial and ethnic groups. They should also examine whether patients who receive an alert can obtain the recommended follow-up. Equal access to a prediction would not, by itself, guarantee equal access to help. These are questions for evaluation, not findings of bias in this study. Medical AI development principles

What would turn a useful signal into better care

Bright’s next evidence milestone would be a prospective evaluation, following new patients as events happen, with a clearly defined care pathway. A strong test would compare care using the score with usual care and track what changes for patients.

The accounting should include time as well as medical outcomes. Who reads the extra result? Who explains uncertainty? Who handles referrals and checks that they happen? A routine service would also need reliable signal checks and a repeatable way to label sleep stages. A sleep clinic should not inherit an open-ended responsibility simply because its equipment collected a useful signal. Medical AI transparency principles emphasize explaining how a system fits into real clinical work. FDA transparency principles

There is a grounded reason for optimism here: researchers have identified a possible new use for information already being collected. The opportunity is to make an existing encounter more helpful. Whether that becomes a benefit people can feel will depend on the care that follows the prediction.

Get Bright Weekly for source-checked AI developments, useful context and the questions that remain open.

Image source: NMCSD Sleep Study Room · DVIDS. Public-use conditions.

How we know7 sources · checked 2026-10-11 · no corrections

Original sources

  1. NHLBI’s guide to sleep studies ↗ · institution
  2. NIH research announcement ↗ · institution
  3. The study in SLEEP ↗ · paper
  4. NHLBI on atrial fibrillation ↗ · institution
  5. NHLBI on heart failure ↗ · institution
  6. Medical AI development principles ↗ · institution
  7. FDA transparency principles ↗ · institution

Institutions: Massachusetts General Hospital · Emory University Hospital · Beth Israel Deaconess Medical Center

Maturity
Emerging
Source published
2026-10-09
Captured
2026-10-11
Last source review
2026-10-11
Editorial method
AI-assisted source review
Place / relevance
Hospital sleep-study cohorts in Boston and Atlanta, United States · unspecified

Bright compared this account with the linked original and supporting sources and kept reported, budgeted, projected, and observed claims distinct. Bright did not independently audit the underlying records.

Maturity describes the tested or operational setting. Confidence describes support for the particular claim; one does not determine the other.

Revision & correction history

2026-10-11T11:37:13Z · Published owner-approved article with retrospective evidence boundaries, exact source links and archival public-domain context photograph.

No corrections recorded.

Keep exploring

Explore the shared question in another setting. These connections do not imply replication.

An editorial connection recorded in Bright’s evidence catalog.Google’s AMIE medical AI interviewed real patients before their appointments ↗Related development · not a replicationGoogle’s AMIE medical AI interviewed real patients before their appointments ↗

Keep looking closer.

See what changed at Bright ↗

Add Bright to your Google Preferred Sources ↗

Suggest a correction · Bright on TikTok