What the Study Found
- An AI grouped nearly 9,300 sleep-study patients into five tiers; the top tier had over double the five-year death risk of the lowest.
- Standard apnea-severity categories showed no significant link to mortality; high-risk patients hid across every severity band.
- The risk tiers predicted death and heart failure in both men and women, unlike the apnea index, which had flagged mainly men.
- The approach held up in a separate, older cohort of over 6,000 people recorded on coarser equipment.
Somewhere in a hospital archive, a night of your breathing is sitting on a hard drive. The rise and fall of your chest, the flicker of your eyes under closed lids, the electrical chatter of your brain drifting down through the sleep stages, your blood oxygen dipping and recovering. All of it recorded, scored once by a technologist, boiled down to a single number, and then more or less left alone. Millions of these nights are gathering dust.
Each year in the US, doctors run somewhere between 1 and 4 million in-lab sleep studies, most of them to check for sleep apnea. And for decades the field has read each one the same narrow way.
The number they pull out is the apnea-hypopnea index (AHI): roughly, how many times an hour your breathing stalls or shallows overnight. It sorts people into mild, moderate and severe, and it decides who gets a breathing machine. But it throws away almost everything else the night recorded. The brain waves, the heart’s rhythm, the shape of the airflow, the way sleep fragments and reassembles: gone, reduced to a tally of pauses. That the AHI leaves something on the table is not a fringe view; a growing body of sleep-medicine research argues the index correlates only modestly with the outcomes that matter and captures little of the heterogeneity among patients, which is why the field has spent recent years hunting for better measures. A team of sleep physicians, data scientists and neuroscientists reckoned there was more in there, and they built an AI to go looking.
Their model, described this week in Nature Communications, learned from more than 10,000 overnight recordings held at the Cleveland Clinic. What it found is a little unsettling, in a useful way.
When the researchers let the model read the full richness of each night and then group patients by what it saw, five distinct clusters fell out. The groups lined up neatly with what happened to those patients over the following years: cardiovascular disease, neurological disease, death. Line the groups up from one to five and the risk climbs at each step. People in the highest-risk group had more than double the mortality risk of those in the lowest over roughly the next five years. This was a retrospective study of nearly 9,300 patients, so it maps associations rather than proving cause, but the separation was stark.
“For decades we have distilled an overnight sleep study into a handful of summary measures,” says Reena Mehra at the University of Washington, the study’s senior clinical author. “AI gives us the opportunity to move beyond those summaries and learn from the full richness of sleep physiology.”
The Apnea Number Missed It
That mortality gap, the one separating the top group from the bottom, was invisible to AHI. When the same patients were sorted the conventional way into mild, moderate and severe apnea, the categories showed no significant link to whether people lived or died. The high-risk patients the model flagged were scattered right across the apnea severity bands, hiding among people the standard test had waved through as fine. The model reads the same raw night; it just refuses to throw the rest of it away. Zeroing out the brain-wave and heart channels scrambled the groupings, which tells you the signal is not just respiratory. It is coming from the whole sleeping body.
There is already a hint in the literature of where that extra signal might live. One of the more promising alternatives to AHI, the sleep apnea-specific hypoxic burden, which measures the depth and duration of overnight oxygen dips rather than just counting pauses, has been shown to predict cardiovascular mortality even after accounting for AHI, in the same community cohorts this study drew on. The AI appears to be reaching for that kind of buried physiological information, only across many channels at once rather than one hand-designed metric.
To build it, the team adapted a transformer, the same broad architecture behind large language models, and pointed it at physiology instead of text. Each 30-second slice of the night became something the model could tokenize and learn from, roughly 126 million internal parameters tuning themselves to the waveforms. The result is a set of numerical fingerprints, one per patient, that carry more prognostic information than any single index the clinic currently prints on its report.
A Number That Travels
A model that works only on its home data is a curiosity, not a tool. So the group took their pipeline to an entirely separate archive, the Sleep Heart Health Study (SHHS), a community cohort recorded years earlier on coarser equipment. The risk groups held up. The highest-risk cluster again tracked with higher mortality and more heart failure, in men and in women alike.
That last detail matters more than it sounds. The old apnea index has a known blind spot: in the original analysis of that same community cohort, sleep-disordered breathing predicted death mainly in men aged 40 to 70, and the researchers couldn’t detect a clear effect in women, having recorded too few deaths among them to be sure. The AI’s groupings didn’t carry that asymmetry. They stratified both sexes, which hints the model is picking up something about risk that the frequency-of-pauses number was quietly missing in women.
“Modern AI lets us recover much more of the information contained in a night’s worth of sleep physiology, revealing clinically meaningful patient groups with very different long-term health risks,” says Jeffrey Rogers at Yale University, the study’s corresponding author.
None of this means the sleep report you already have is about to change. The study is retrospective, built from records rather than a trial, and it can’t say whether acting on these risk groups would actually help anyone. The team could not see how faithfully patients used their breathing machines, medication histories were patchy, and residual confounding is always lurking in this kind of data. There’s a subtler catch too. The training nights all came from a single US health system, however diverse its patients, and the production code stays locked away behind the industrial partner’s intellectual-property controls, which makes independent replication harder than anyone would like.
Still, the wider claim lands. Routine tests, ordered by the million and read for one thing, may be quietly carrying far more about our health than we bother to extract.
- Study type: Retrospective clinical cohort with a transformer-based AI (foundation) model and external validation; peer-reviewed, published in Nature Communications
- Sample size: 9,608 overnight sleep recordings from 9,297 patients (Cleveland Clinic, 2012โ2022); external cohort of over 6,000 (Sleep Heart Health Study)
- Exposure: AI-derived physiologic risk groups (five clusters) from full polysomnography signals
- Comparison group: Conventional apneaโhypopnea index severity categories (mild, moderate, severe)
- Follow-up: Mean 4.9 years for mortality; total median observation window about 15 years
- Funding / conflicts of interest: Cleveland ClinicโIBM Discovery Accelerator Program and the US National Heart, Lung, and Blood Institute; authors declare no competing interests, though two are IBM Research employees
- Data availability: Sleep Heart Health Study data available via SleepData.org; Cleveland Clinic data restricted under IRB and data-use agreement; production model code withheld citing IBM intellectual-property controls
- Main limitation: Retrospective design limits causal inference, and breathing-machine (PAP) adherence and medication data were unavailable, so residual confounding cannot be excluded
Reference
Bilal, E., Araujo, M. L. D., Beck, K. L., Heinzinger, C. M., Ghosn, S., Saab, C. Y., Foldvary-Schaefer, N., Rogers, J. L., & Mehra, R. (2026). A foundation model for sleep-based risk stratification and clinical outcomes. Nature Communications, 17(1). https://doi.org/10.1038/s41467-026-75326-9
Frequently Asked Questions
Why does an AI reading old sleep tests matter for people who already had a normal apnea result?
An AI reading old sleep tests matters for people with a normal apnea result because the model flagged high-risk patients scattered across every apnea severity band, including some the standard index had cleared. The apnea-hypopnea index counts breathing pauses and little else, so it can miss risk written into the brain waves, heart rhythm and sleep structure that the AI reads from the same recording. Whether acting on that risk helps anyone is still untested.
Does this study prove that disrupted sleep causes heart disease or death?
No, this study does not prove that disrupted sleep causes heart disease or death. It is a retrospective analysis of records from nearly 9,300 patients, so it can show that the AI’s risk groups are strongly associated with later cardiovascular disease, neurological disease and mortality, but it cannot establish cause. The authors are explicit that the retrospective design limits any causal conclusion.
How does the model actually pull more out of a sleep study than the standard number?
The model pulls more out of a sleep study by using a transformer, the same broad kind of AI behind large language models, to read the full overnight recording rather than a single summary score. It learns numerical fingerprints from the brain, heart, breathing and oxygen signals across the night, then groups patients by those fingerprints into five risk tiers. When the researchers blanked the brain-wave or heart channels, the groupings fell apart, showing the signal is not just respiratory.
Could this approach be used on other routine medical tests, not just sleep studies?
This approach could in principle be used on other routine medical tests, and that is the study’s broader implication: everyday tests ordered by the million may carry far more health information than clinicians currently extract from them. Sleep was simply a convenient starting point, since nearly everyone sleeps and the recordings already exist in vast numbers. Proving the idea elsewhere would need the same kind of validation this model is still working through.
What’s stopping this from being used in clinics tomorrow?
What’s stopping this from being used in clinics tomorrow is that it has only been validated on existing records, not tested in a trial where doctors act on the risk groups. The training data came from a single US health system, key details like breathing-machine adherence were missing, and the production code is held back under the industrial partner’s intellectual-property rules, all of which make independent replication and real-world deployment harder. Prospective studies are the necessary next step.
Cite This Page
