BlogSleep16 min read

How accurate is Apple Watch sleep tracking

Apple Watch sleep stage accuracy against polysomnography, how it compares with Oura and WHOOP, and why a night can go missing.

Published September 25, 2026
This content is for informational purposes only and is not a substitute for professional advice.

In a Belgian sleep lab, an Apple Watch Series 8 agreed with polysomnography on sleep stages at a Cohen's kappa of 0.53, the best of six wrist devices worn by the same 62 adults (Schyvens et al., 2025). In a second lab study, the same model scored a kappa of 0.30 and a macro F1 of 0.49, the lowest of five wearables tested (Lee et al., 2023). Both studies scored the watch against polysomnography, the clinical reference, and both tested the watchOS 9 algorithm that first brought sleep stages to the watch. Read together, the two results show what the watch can and cannot tell you about a night better than either one alone.

This article puts numbers where most consumer coverage hedges, corrects two studies that are routinely miscited as Apple Watch evidence, and covers the settings and Titan rules that decide whether a night appears at all. For the general mechanics of consumer sleep tracking, start with the glossary entry.

01How Apple Watch detects a night of sleep

Polysomnography records brain waves, eye movements and chin muscle tone all night, and a technologist labels every 30 second epoch under American Academy of Sleep Medicine rules as wake, REM, or one of three non-REM stages called N1, N2 and N3. A watch has no electrodes on your head. Every stage it reports is an inference from the wrist.

Apple's staging classifier runs on the accelerometer. Apple's technical paper, last updated in October 2025, describes a model that reads high-frequency 3-axis accelerometer data and assigns each 30 second epoch to Awake, REM, Core or Deep (Apple, 2025). At that sampling rate it picks up the small movements of breathing along with gross body movement, and breathing rate and regularity shift from stage to stage. Much consumer coverage says the watch stages sleep from heart rate. Apple's documentation for the staging model names accelerometer data only. The watch records heart rate, respiratory rate and blood oxygen overnight as separate measurements.

Apple's labels map onto clinical stages. Deep is N3, slow-wave sleep. Core is N1 plus N2, the stage most other brands call light sleep. Apple avoided the word "light" because N2 usually fills more than half of a healthy night and carries the sleep spindles and K-complexes that sleep researchers treat as markers of normal sleep. Awake and REM keep their clinical meanings.

What changed with watchOS 9, 11 and 26

Sleep tracking arrived in 2020 with watchOS 7, and it classified each 30 second window as asleep or awake. Stages arrived in 2022 with watchOS 9. In 2024, watchOS 11 began detecting sleep outside a set schedule, including naps. For a detected session with one to three hours of sleep, the watch writes plain asleep and awake data to Apple Health. It writes full stages only when a session passes three hours of sleep. In 2025, watchOS 26 retrained the staging model with foundation models built from Apple Heart and Movement Study data. Apple aimed the update at quiet wake, the stretches before and after sleep when you lie awake and still.

On its original validation set, Apple reports that the watchOS 26 model raised correct wake classification from 70 to 79 percent and median kappa from 0.63 to 0.68 (Apple, 2025). Every independent study in this article tested the watchOS 9 era algorithm, and no independent lab has yet published a polysomnography test of the 2025 model. De Zambotti and colleagues (2025) list this version lag among the field's standing problems. Validation papers routinely reach print after the firmware they tested has been replaced.

02Two studies cited for a device they did not test

Two papers appear in most articles about Apple Watch sleep accuracy. Neither measured the feature Apple ships today.

Chinoy and colleagues (2021) ran a careful polysomnography study in 34 healthy young adults over three lab nights, one of them deliberately fragmented with auditory tones. They tested seven devices, the Fatigue Science Readiband, Fitbit Alta HR, Garmin Fenix 5S, Garmin Vivosmart 3, EarlySense Live, ResMed S+ and SleepScore Max. Apple Watch was not one of them. The paper is a strong source on the wake detection problem that every wrist device shares. It contains no Apple Watch data.

Stone and colleagues (2020) did include an Apple Watch. It was a Series 3 running two third-party apps, SleepWatch and Sleep++, the usual route to watch sleep data before Apple shipped its own tracking in watchOS 7. Neither app reported sleep stages. The reference was the Sleep Profiler, an at-home EEG headband, across 98 nights from five adults. Sleep++ overestimated total sleep by 0.80 hours and SleepWatch by 1.17 hours. Those figures describe two third-party apps on 2017 hardware and say nothing about the staging model Apple released in 2022. The paper's conclusion that no device it tested quantified sleep stages accurately rests on the devices that did report stages, which included WHOOP Strap 2.0 and the second-generation Oura ring.

03Total sleep time and the wake problem

Every wrist device in the validation literature finds sleep well and finds wakefulness poorly. Chinoy (2021) reported sleep sensitivity of 0.93 to 0.99 across its seven devices and wake specificity of 0.18 to 0.54. The asymmetry comes from the sensor. Lying awake and still produces almost the same wrist signal as sleep, so most errors run one way, and the device credits you with sleep you did not get.

Apple Watch follows the pattern with better wake numbers than most. In Schyvens (2025) it detected 96.3 percent of true sleep epochs and 52.2 percent of true wake epochs, the best wake figure of the six devices. It overestimated total sleep time by 19.6 minutes and underestimated wake after sleep onset by 21.2 minutes, and both differences were statistically significant. Apple's own validation set, 166 participants and 299 nights held out from training, reports a median sleep sensitivity of 97.9 percent and specificity of 75.0 percent for the watchOS 9 model, and 96.8 and 78.9 percent after the watchOS 26 update. Apple reports per-night medians from its own study, while the independent labs pool epochs, and the published data cannot say how much of the gap comes from method and how much from who ran the study.

A meta-analysis of 24 studies and 798 participants found wrist trackers as a class missed total sleep by about 17 minutes on average (Lee et al., 2025). It included a single Apple Watch wearer, so it describes the category. Taken with Schyvens, it puts the average total sleep error for a current wrist device near 20 minutes a night, with single nights spread wider.

04How well Apple Watch separates REM, deep and core sleep

Stage agreement is where the numbers fall. Lambe and colleagues (2026), in a living systematic review of Apple Watch accuracy, pooled three sleep staging studies with 221 participants, one of them Apple's own validation. They found good separation of sleep from wake, with two studies reporting sleep sensitivity of 97 percent or higher, and moderate to poor separation of the sleep stages from one another. One of the included studies, in healthy adults wearing a Series 8, found that the watch significantly underestimated deep sleep and overestimated lighter sleep.

Stage by stage, Schyvens (2025) found the Series 8 labeled 83.3 percent of true Core epochs correctly, 68.6 percent of REM, 50.7 percent of Deep and 52.2 percent of wake. It underestimated deep sleep by 25.2 minutes a night. Apple's confusion matrix for the watchOS 9 model shows 83 percent of Core, 78 percent of REM, 62 percent of Deep and 70 percent of wake classified correctly, with 38 percent of true deep sleep called Core. For the watchOS 26 model Apple reports 82, 82, 68 and 79 percent, and the share of deep sleep called Core falls to 32 percent.

The table lines up the independent studies and Apple's own figures. Every row comes from a separate study with its own sample, protocol and statistic. Read it for the range, and do not average down a column.

Device and algorithmStudyParticipantsTrue sleep detectedTrue wake detectedStage agreementTotal sleep bias
Apple Watch, watchOS 9 modelApple 2025, manufacturer166 (299 nights)97.9%75.0%Kappa 0.63Not reported
Apple Watch, watchOS 26 modelApple 2025, manufacturer166 (299 nights)96.8%78.9%Kappa 0.68Not reported
Apple Watch Series 8Schyvens 20256296.3%52.2%Kappa 0.53+19.6 min
Apple Watch 8Lee 202326Not reported44.8%Kappa 0.30, macro F1 0.49Not reported
WHOOP 4.0Schyvens 20256293.6%40.1%Kappa 0.37+24.5 min
WHOOP, pooledSchyvens 2024, four studiesPooled91.7%55.7%Kappa 0.46-1.4 min
Oura, first-generation algorithmde Zambotti 20194196%48%51% deep and 61% REM agreement-1.3 min
Fitbit SenseSchyvens 20256293.3%48.8%Kappa 0.42+6.3 min
Garmin Vivosmart 4Schyvens 20256295.9%29.4%Kappa 0.21+38.4 min

Best of six in one lab, worst of five in another

Schyvens (2025) and Lee (2023) disagree by more than the gap between most devices in either study. Three design differences point toward the cause, and none of them makes either study wrong.

The metrics differ. Macro F1 averages the four stage-level F1 scores with equal weight, so the watch's weak Deep score in Lee (F1 0.31) counts as much as its Core score (0.67), even though Core makes up most of the night. Kappa is computed over every epoch, where Core dominates. The metric alone cannot explain the result, because the two kappas, 0.53 and 0.30, also disagree.

The alignment differs. Schyvens split each wearable epoch into 30 second pieces to match the polysomnography record. Lee cut every device's output and the reference into 1 second slices and compared them second by second. De Zambotti and colleagues (2025) warn, as a general point about the literature, that comparing a device at a resolution its algorithm was not built for creates artificial stage transitions and lowers measured accuracy.

The samples differ. Lee recruited adults with sleep complaints from a tertiary hospital and a sleep clinic, and 26 of its 75 participants wore an Apple Watch. Schyvens tested 62 adults, 52 of them men, with a mean age of 46, a mix of people referred for suspected sleep apnea and healthy volunteers. Apple's own data shows why that matters. In a held-out clinical cohort of 236 patients undergoing diagnostic polysomnography, kappa fell to 0.55 for the watchOS 9 model, and Apple's analysis found that frequent stage transitions and a higher apnea-hypopnea index lowered agreement more than age or sex did. Both independent studies also recorded a single night per person.

The rankings in Lee also sat close together. The five wearables scored macro F1 values between 0.49 and 0.58, with Fitbit Sense 2 at 0.58, Galaxy Watch 5 at 0.58, Pixel Watch at 0.57, Oura Ring 3 at 0.52 and Apple Watch 8 at 0.49. "Worst of five" describes a 0.09 spread. In Schyvens the six devices spread from 0.21 to 0.53 in kappa, a wider field in which Apple Watch sat clearly at the top.

The independent evidence puts Apple Watch staging in the fair to moderate range, near the top of current wrist devices, with too much study-to-study variation for one ranking to decide a purchase. Apple states in its own paper that watch sleep stages are not intended for clinical use, and no study here gives a reason to use them that way.

Why deep sleep is where every device misses

Deep sleep is a small share of the night. In Apple's validation set it averaged 12.8 percent of sleep, against 65.9 percent Core and 21.4 percent REM. A classifier that is unsure about an epoch does best on average by calling the majority class, and here the majority class is Core. The wrist signals also overlap. N2 and N3 are both periods of stillness with slow, regular breathing, and the difference between them lies mainly in brain wave amplitude that a wrist sensor cannot see. REM, with irregular breathing and near-total loss of muscle tone, looks more distinct. In Apple's data the watch called only 0.13 percent of true deep sleep wake and 0.28 percent of true REM deep.

The pattern holds across brands. The first Oura algorithm agreed with polysomnography on 51 percent of deep sleep and underestimated it by about 20 minutes (de Zambotti et al., 2019). The WHOOP studies pooled by Schyvens (2024) underestimated deep sleep by 9.3 minutes on average. Across every device Stone (2020) tested, the mean absolute percentage error for deep sleep was 67.96 percent. A low deep sleep number is the most likely error on any consumer tracker.

05The errors you will see on your own sleep report

The literature and Apple's confusion matrices point to five recurring errors, each with a recognizable look on a sleep timeline.

Quiet wake counted as sleep. Reading in bed with the lights low, or lying still after the alarm, can land in Core. The watchOS 9 model labeled 27 percent of true wake epochs as Core, and Apple built the watchOS 26 update around this error.

Short awakenings that vanish. Brief wake-ups in the middle of the night are the epochs a still wrist hides best. Schyvens measured a 21.2 minute underestimate of wake after sleep onset, so a night that felt broken can show as unbroken.

Deep sleep that looks low. A third or more of true deep sleep lands in Core. One low deep night is more likely a classification miss than a change in your physiology.

REM and Core trading places. The watchOS 9 model labeled 21 percent of true REM as Core. After the watchOS 26 update Apple reports 16 percent.

Naps and split nights handled differently. A nap of one to three hours reaches Apple Health as plain asleep time with no stages. A night broken by a long wake can reach Titan as two separate sessions, covered below.

06How Apple Watch compares with Oura and WHOOP

No published study has put Apple Watch, Oura and WHOOP on the same people on the same night against polysomnography. Schyvens (2025) tested Apple Watch and WHOOP 4.0 without Oura. Lee (2023) tested Apple Watch and Oura Ring 3 without WHOOP. Any three-way comparison is a synthesis across separate studies.

Within each pairing, the results line up. Against WHOOP 4.0 in Schyvens, Apple Watch had higher kappa (0.53 against 0.37), detected more true wake (52.2 against 40.1 percent) and overestimated total sleep by less (19.6 against 24.5 minutes). WHOOP found more of the true deep sleep, 69.6 percent against 50.7. Against Oura Ring 3 in Lee, the ring scored a slightly higher macro F1 (0.52 against 0.49) and found far more true deep sleep, with a deep sensitivity of 77.8 percent against 41.3. In both pairings the competitor found more deep sleep and Apple Watch undercalled it. The pooled WHOOP kappa of 0.46 (Schyvens et al., 2024) sits in the same moderate band. All three devices share the same limits on different hardware, and none of them wins on every measure.

In Titan, the route the data takes matters as much as the device. Oura sleep that the Oura app writes to Apple Health carries Core, Deep and REM stages. Oura sleep from Titan's direct Oura connection arrives as one asleep block with no stage breakdown, as the Oura support article explains. WHOOP writes sleep with duration and stages to Apple Health, and it writes no HRV. A WHOOP-only setup therefore gets a staged Sleep score and no Recovery, Stress or Battery score, as the WHOOP support article lists. For how each brand's other metrics hold up, see the wearables data quality guide and the glossary entry on wearable devices.

07Settings that decide whether a night records at all

Accuracy only matters once a night exists. Five conditions decide whether a staged night reaches Titan.

Sleep tracking has to be on. In the Health app, go to Browse > Sleep > Full Schedule & Options and turn on Track Sleep with Apple Watch, as the Apple Watch support article describes. The watch has to be on your wrist in bed, since it records sleep stages, respiratory rate, blood oxygen and sleep HRV only while you wear it. It needs enough charge to last the night. Apple prompts you to charge if the battery is below 30 percent at bedtime (Apple Support). Fit counts too. Wear that departs from the manufacturer's recommendation measurably changes algorithm performance (de Zambotti et al., 2025), so wear the watch snug.

The last condition sits on the iPhone. Titan reads sleep from Apple Health, so Sleep must be allowed under Titan's Health permissions. When Titan cannot read sleep, it shows no sleep and no error.

08Why a night can be missing from Titan

A night the watch never recorded cannot be recovered downstream. A night it did record can still drop out of a day under Titan's session rules, which how sleep is detected sets out in full.

Titan sorts every sleep sample in Apple Health by start time. A sample that starts within 2 hours of the latest end time so far joins the current session, and a longer gap starts a new one. Titan then assigns each session to the day it ends on, in your phone's local time, provided it ends after midnight and no later than 12:00 PM. Sleep that ends after noon counts for no day. A lie-in until 12:30 PM, a daytime sleep after a night shift and an afternoon nap all fall outside the window.

When more than one session ends on the same morning, the session with the most time asleep wins. Time asleep is the sum of Core, Deep, REM and Asleep time, and Awake and In Bed time do not count. Say you sleep from 11:00 PM to 2:30 AM, lie awake with no samples until 5:15 AM, and sleep again until 8:00 AM. The gap is 2 hours 45 minutes, so Titan sees two sessions. It keeps the 3.5 hour first half and drops the 2 hours 45 minutes after it, and the dropped sleep adds nothing to that day's total.

When you wear two trackers, Titan cuts the night into slices at every point where a sample starts or ends and keeps the highest-priority stage for each slice. Awake beats REM, REM beats Deep, Deep beats Core, and Core beats plain Asleep. Since a still wrist often misses a wake-up, that order keeps any wake-up either tracker catches, and a staged tracker overrides a duration-only one for the same minutes. The sleep missing or from the wrong source article works through five checks in order, ending with the "Sleep tracked by" line on the Sleep card.

09What to trust on your report

Given moderate stage agreement in every independent study, the reliable numbers on an Apple Watch sleep report are total sleep time and its consistency from night to night. Single-night REM and Deep minutes deserve the least weight. A deep sleep figure that drops 25 minutes on one night sits inside the watch's measured error. The same drop held across two weeks on the same watch is a trend worth attention, since comparing a device with itself cancels whatever part of its error repeats every night.

Titan's Sleep score gives the largest share to the most reliable measure. Duration against your sleep target carries 44 percent of the score. Restorative sleep, Deep plus REM as a share of time asleep with a bonus for an even split between them, carries 35 percent. Awake time as a share of time in bed carries 21 percent. Titan treats 40 percent Deep plus REM as a typical night and gives full credit at that share. On a night with 7 hours asleep, the 25 minute deep sleep shortfall Schyvens measured moves the Deep plus REM share by about 6 percentage points. The watch's tendency to miss short awakenings pushes the other way, since less recorded wake means more credit on the awake time part.

Nights without stages score on duration alone. Meeting your target on a duration-only night scores 100, and nothing counts against it for low Deep and REM or for time awake, so a duration-only night can outscore a staged night of the same length. If your scores jump after you change trackers, or after you move Oura sleep to the direct connection, check whether the new source records stages. For keeping timing steady, the sleep hygiene entry and Titan's sleep consistency article cover the habits and the measure.

10The Sleep Tracker setting in Titan

Night-to-night consistency is the part of the data you can trust, and Sleep Tracker is the setting that keeps Titan reading the same device every night. Open Settings > Recovery & Sleep > Sleep Preferences > Sleep Tracker. You reach Settings from the avatar on Today, then You > Settings. The default is Apple Watch, and it matches any Apple Health sleep source whose name contains "watch" or "apple".

Titan builds each day from your chosen tracker first. On a day that tracker has no session, Titan falls back to all your sleep sources, so a night you left the watch on the charger can still use sleep from a ring. The HealthKit Sources entry for Sleep, under Settings > Health & Connected Apps > Data Sources & Integrations, overrides Sleep Tracker whenever you set a primary or allowed source there, and it has no fallback. Sleep Tracker has no WHOOP option, so a WHOOP wearer sets WHOOP as the primary source under HealthKit Sources > Sleep instead. With Oura as the Sleep Tracker, Titan matches only the sleep the Oura app writes, and it skips direct-connection records.

If Apple Watch is the only device on your wrist at night, leave Sleep Tracker on Apple Watch. If you wear two trackers, pick the one you wear most consistently and let the other fill the gaps. The choosing your sleep source article covers every combination.

11References

  • Apple Inc. (2025). Estimating sleep stages from Apple Watch, updated October 2025. Apple technical paper. https://www.apple.com/health/pdf/EstimatingSleepStagesfromAppleWatchOct_2025.pdf
  • Apple Support. Track your sleep on Apple Watch. Apple Watch User Guide. https://support.apple.com/guide/watch/track-your-sleep-apd830528336/watchos
  • Chinoy ED et al. (2021). Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep. https://doi.org/10.1093/sleep/zsaa291
  • de Zambotti M et al. (2019). The sleep of the ring: comparison of the ŌURA sleep tracker against polysomnography. Behavioral Sleep Medicine. https://doi.org/10.1080/15402002.2017.1300587
  • de Zambotti M et al. (2025). Toward better evaluation of consumer sleep technologies: a call for rigor, context, and collaboration. Sleep Advances. https://doi.org/10.1093/sleepadvances/zpaf063
  • Lambe R et al. (2026). The accuracy of Apple Watch measurements: living systematic review and meta-analysis. NPJ Digital Medicine. https://doi.org/10.1038/s41746-025-02238-1
  • Lee T et al. (2023). Accuracy of 11 wearable, nearable, and airable consumer sleep trackers: prospective multicenter validation study. JMIR mHealth and uHealth. https://doi.org/10.2196/50983
  • Lee YJ et al. (2025). Performance of consumer wrist-worn sleep tracking devices compared to polysomnography: a meta-analysis. Journal of Clinical Sleep Medicine. https://doi.org/10.5664/jcsm.11460
  • Schyvens AM et al. (2024). Accuracy of Fitbit Charge 4, Garmin Vivosmart 4, and WHOOP versus polysomnography: systematic review. JMIR mHealth and uHealth. https://doi.org/10.2196/52192
  • Schyvens AM et al. (2025). A performance validation of six commercial wrist-worn wearable sleep-tracking devices for sleep stage scoring compared to polysomnography. Sleep Advances. https://doi.org/10.1093/sleepadvances/zpaf021
  • Stone JD et al. (2020). Evaluations of commercial sleep technologies for objective monitoring during routine sleeping conditions. Nature and Science of Sleep. https://doi.org/10.2147/NSS.S270705
Keep readingAll stories