AI finds heart risk clues in overnight sleep tests
A new study suggests sleep-test recordings could help flag future heart problems. The next challenge is proving that the extra warning improves care.
- Maturity
- Emerging, stage 2 of 4
- Support
- 7 sources · institution, paper
- Evidence detail
- How we know ↓

An overnight sleep test gathers hours of information about someone’s breathing, brain activity and heartbeat. Doctors use it to investigate problems such as sleep apnea, when breathing repeatedly stops and starts. That night of monitoring may also contain clues to a patient’s longer-term health. NHLBI’s guide to sleep studies
Research highlighted by the National Institutes of Health on October 9 explores one possibility: using artificial intelligence to extract information about future heart risk from the ECG, a recording of the heart’s electrical activity, already collected during a sleep study. The system combines that signal with information about sleep stages, the phases a person moves through while sleeping. NIH research announcement
The attraction is practical. If the approach eventually proves useful, patients undergoing these tests could get more value from the same night of monitoring. For now, the evidence comes from past records. The study does not establish that using its predictions improves patients’ health. The study in SLEEP
What the researchers found
The model was fine-tuned using 15,809 patients at Massachusetts General Hospital. Researchers then tested it on separate groups: 9,810 patients at Emory University Hospital and 12,576 at Beth Israel Deaconess Medical Center. It used a single ECG channel and expert-labeled sleep stages. NIH research announcement
The study targeted outcomes over 10 years. Adding the model’s output to clinical and sleep-related risk factors improved discrimination for atrial fibrillation, heart failure and death from any cause. Discrimination means distinguishing people who later experience an outcome from those who do not. The model did not improve that distinction for stroke or heart attack. The study in SLEEP
Atrial fibrillation is an irregular heart rhythm. Heart failure means the heart cannot pump enough blood to meet the body’s needs. Both are serious conditions, but a prediction of greater future risk does not tell a person that they have either condition now. NHLBI on atrial fibrillation, NHLBI on heart failure
Why testing at other hospitals matters
The separate hospitals are an important part of the result. A model can appear impressive when it learns patterns peculiar to the place that supplied its training data. Testing on patients from other hospitals asks a tougher question: does the useful signal survive a change of setting?
Here, the encouraging interpretation is that the method deserves investigation beyond its original development site. It still needs testing across the range of people, equipment and working conditions where clinicians might use it. General principles for medical AI explicitly call for representative patients and independent test data. Medical AI development principles
The study population consisted of hospital sleep-study patients. It does not establish performance for everyone who wears a watch to bed or uses a home sleep device. The study in SLEEP
That boundary matters to the promise. A useful tool for a defined group can still be valuable. Reaching people outside that group is a separate research question.
A warning needs to mean something
Imagine a future clinic receiving a risk score alongside a sleep report. Before acting, staff would need to know how accurately the number describes patients like theirs. If a system assigns a 10 percent risk to a group, roughly one in ten should experience the specified outcome within the stated period. This is calibration, a different question from whether the system ranks people in the right order.
The paper reports calibration over a six-year follow-up period and retrospective decision-curve analyses. Those checks do not establish the effects of using the model in care. The study in SLEEP
For this application, a useful next study would specify an actual decision: which scores trigger a review, who conducts it and what information they consider. It should measure how many people receive worthwhile follow-up and how many are sent through unnecessary appointments. A mathematically better prediction could still be difficult to use if the next step is vague.
The limits of looking backward
The researchers identified cardiovascular diagnoses through medical-record codes, without direct clinical confirmation. These records can contain errors. They also acknowledged the difficulty of separating future atrial fibrillation from previously unrecognized episodes. The study in SLEEP
That raises a practical question for further testing: what, exactly, is the score picking up in each patient? A signal of existing illness and a warning of future illness can lead to different conversations and different investigations. Clinicians need enough information to interpret the result responsibly.
Fairness deserves similarly specific scrutiny. Future evaluations should ask whether errors concentrate in particular age groups, sexes or racial and ethnic groups. They should also examine whether patients who receive an alert can obtain the recommended follow-up. Equal access to a prediction would not, by itself, guarantee equal access to help. These are questions for evaluation, not findings of bias in this study. Medical AI development principles
What would turn a useful signal into better care
Bright’s next evidence milestone would be a prospective evaluation, following new patients as events happen, with a clearly defined care pathway. A strong test would compare care using the score with usual care and track what changes for patients.
The accounting should include time as well as medical outcomes. Who reads the extra result? Who explains uncertainty? Who handles referrals and checks that they happen? A routine service would also need reliable signal checks and a repeatable way to label sleep stages. A sleep clinic should not inherit an open-ended responsibility simply because its equipment collected a useful signal. Medical AI transparency principles emphasize explaining how a system fits into real clinical work. FDA transparency principles
There is a grounded reason for optimism here: researchers have identified a possible new use for information already being collected. The opportunity is to make an existing encounter more helpful. Whether that becomes a benefit people can feel will depend on the care that follows the prediction.
Get Bright Weekly for source-checked AI developments, useful context and the questions that remain open.
Image source: NMCSD Sleep Study Room · DVIDS. Public-use conditions.
How we know7 sources · checked 2026-10-11 · no corrections
Original sources
- NHLBI’s guide to sleep studies ↗ · institution
- NIH research announcement ↗ · institution
- The study in SLEEP ↗ · paper
- NHLBI on atrial fibrillation ↗ · institution
- NHLBI on heart failure ↗ · institution
- Medical AI development principles ↗ · institution
- FDA transparency principles ↗ · institution
Institutions: Massachusetts General Hospital · Emory University Hospital · Beth Israel Deaconess Medical Center
- Maturity
- Emerging
- Source published
- 2026-10-09
- Captured
- 2026-10-11
- Last source review
- 2026-10-11
- Editorial method
- AI-assisted source review
- Place / relevance
- Hospital sleep-study cohorts in Boston and Atlanta, United States · unspecified
Bright compared this account with the linked original and supporting sources and kept reported, budgeted, projected, and observed claims distinct. Bright did not independently audit the underlying records.
Maturity describes the tested or operational setting. Confidence describes support for the particular claim; one does not determine the other.
Revision & correction history
2026-10-11T11:37:13Z · Published owner-approved article with retrospective evidence boundaries, exact source links and archival public-domain context photograph.
No corrections recorded.
Keep exploring
Explore the shared question in another setting. These connections do not imply replication.
