A baseline is your own typical value for a body measurement, such as overnight HRV or sleeping heart rate, calculated from your recent history together with how much that value normally varies from day to day. A wearable compares each new reading with your baseline to decide whether today is unusual for you.
Most recovery, readiness and stress scores on watches and rings compare the latest reading with a baseline and pass the difference through a scoring curve. The baseline decides what counts as normal, so it decides what the score means. Two apps that read the same HRV can disagree about your morning because one looks back 7 days and the other 60, or because one averages and the other takes a median. Once you know how a baseline is built, you can tell when a low score reflects your body and when it reflects the arithmetic. The wearable data quality guide covers how accurate each input is before it reaches a baseline.
01Why your own history is the right reference
People differ from one another far more than one person differs from day to day. Altini and Plews analyzed about 9 million morning measurements from 28,175 users of the HRV4Training app (Sensors, 2021). Mean RMSSD across the dataset was 69 ± 37 ms. Age, sex, body mass index and activity level together explained only 15 percent of the variance in RMSSD and 19 percent in resting heart rate. A chart built from those four facts leaves most of the differences between people unexplained.
The changes you care about day to day are small by comparison. In the same data, HRV fell 3 percent after high-intensity training, about 10 percent during sickness and 12 percent after heavy drinking. A 12 percent drop in a person with an HRV of 90 ms still leaves them far above most of a population chart. The same drop in a person at 30 ms may look alarming on the chart and still be ordinary for them. The chart cannot see either change. Your own history can.
Clinical chemistry formalized this problem decades ago. Fraser and Harris (1989) described how to split a lab value's variation into within-person and between-person parts, and how that split decides whether population reference values are useful for an individual patient. The ratio of the two is called the index of individuality. Wu and colleagues (2025), applying it to ECG intervals, state the rule that labs use. "An index of less than 0.6 indicates that a population-based reference interval is of no value," and serial change becomes the useful measure. Nobody has published a formal index for wearable HRV. A rough ratio from two separate sources, a day-to-day error of about 12 percent for log RMSSD (Buchheit, 2014) against a between-person spread of 37 ms on a 69 ms mean (Altini and Plews, 2021), comes out near 0.2. That points the same way. Your own history is the reference that can register a 10 percent drop.
02The two parts of a baseline
Every baseline needs a center, meaning what your number usually is, and a spread, meaning how far it usually moves. The center alone tells you that tonight is 4 ms low. The spread tells you whether 4 ms is a normal wobble or an unusual night for you.
| Method | Center or spread | Strength | Weakness |
|---|---|---|---|
| Rolling mean | Center | Simple, smooths daily noise | One bad reading moves it |
| Median | Center | Ignores a few extreme readings | Holds its old value until about half the window has shifted |
| Standard deviation | Spread | Standard in research | One bad reading inflates it |
| Median absolute deviation | Spread | Ignores a few extreme readings | Needs scaling by 1.4826 to match a standard deviation |
| Coefficient of variation | Spread, as a percentage | Compares noise across people and measures | Depends on the mean being stable |
| Smallest worthwhile change | Decision threshold | Ties the band to a meaningful change | No agreed size for HRV |
Rolling means and log-transformed HRV
The athlete-monitoring literature built its methods on the rolling mean. Plews and colleagues followed two elite triathletes for 77 days (European Journal of Applied Physiology, 2012). They tracked the 7-day rolling average of log-transformed RMSSD, written Ln rMSSD. In the athlete who became non-functionally overreached, that average fell steadily in the weeks before a race where the athlete performed poorly. The control athlete's average stayed flat. HRV values are right-skewed, so researchers take the natural log first, which turns a percentage change into a fixed step on the scale.
Averaging matters more than the choice of metric. In 10 runners over a 9-week program, Plews and colleagues (International Journal of Sports Physiology and Performance, 2013) found that the change in maximal aerobic speed had a trivial correlation with a single day's Ln rMSSD (r = −0.06) and a very large correlation with the weekly average (r = 0.72). A later analysis by the same group found the benefit of averaging levels off after 3 to 4 days, and recommended at least 3 valid readings per week (Plews et al., 2014).
Median and median absolute deviation
A mean and a standard deviation both chase outliers. Leys and colleagues (2013) note that the mean has a breakdown point of 0, because a single infinite value makes it infinite, while the median has a breakdown point of 50 percent. The median absolute deviation (MAD) gives the same protection for spread. You take the median, measure each value's absolute distance from it, and take the median of those distances. Multiplying by 1.4826 puts it on the same scale as a standard deviation when the data are normally distributed.
Here is how that plays out with 15 nights of HRV between 38 and 52 ms and a median of 45 ms. The mean is 45.3 ms and the standard deviation 4.3 ms, so a 40 ms night sits 1.25 standard deviations low. Now replace one night with a 140 ms reading from a loose band. The mean jumps to 51.3 ms, the standard deviation to 24.9 ms, and the same 40 ms night now looks only 0.45 standard deviations low. The median and MAD do not move at all. One faulty reading was enough to hide a bad night from the mean-based baseline. The median-based baseline still flags it.
Coefficient of variation
The coefficient of variation (CV) is the standard deviation divided by the mean, expressed as a percentage. Hopkins (2000) recommends it for expressing typical error, the standard deviation of one person's repeated measurements, because a percentage carries across people with very different absolute values. Buchheit (2014) puts the typical error at about 12 percent for resting Ln rMSSD and about 10 percent for resting heart rate. Al Haddad and colleagues (2011) found CVs of 4 to 17 percent for time-domain HRV indices across repeated seated recordings.
The CV can also be a signal in its own right. In the Plews 2012 case study, the CV of the overreached athlete's 7-day Ln rMSSD fell by 0.65 percent per week on the way to the failed race. That athlete's HRV became less variable as well as lower.
Smallest worthwhile change and the normal-range approach
The smallest worthwhile change (SWC) is the smallest shift that matters in practice. Buchheit's review lists about +3 percent for resting vagal HRV and about −2 percent for resting heart rate as the changes that go with real improvements in fitness. Both are smaller than the day-to-day typical error, which gives signal-to-noise ratios of 0.8 and 0.7. A single morning cannot reliably show a change that small. An average over several days can.
For individuals, the Plews and Buchheit approach compares a rolling average with a band around the athlete's baseline, with the band's width set by the SWC. Buchheit notes that authors have used 0.5 × CV (Le Meur and colleagues) or 1 × CV (Plews and colleagues) as the SWC. He adds a warning that applies to every wearable. "There is currently no evidence that changes greater than any fraction of the CV would actually be meaningful in practice." The band width is a judgment call, and it should stay the same over time.
Training studies built on this idea give it practical support. Kiviniemi and colleagues (2007) set each runner's reference as the 10-day mean minus one standard deviation. A reading below it, or a falling trend for 2 days, meant easy training or rest. Over 4 weeks the HRV-guided group raised VO2peak from 56 to 60 ml/kg/min, while the fixed-plan group showed no significant change. Vesterinen and colleagues (2016) gave 40 recreational runners 4 weeks of preparation to establish their values, then prescribed moderate and hard sessions only when HRV sat inside an individual SWC band. The guided group did fewer hard sessions (13.2 against 17.7) and improved 3,000 m time by 2.1 percent. The fixed-plan group's 1.1 percent change was not significant. Javaloyes and colleagues (2019) used 4 baseline weeks in well-trained cyclists before starting HRV-guided training. The HRV and training readiness guide turns these rules into a daily routine.
03How long the window should be
Window length decides what the baseline treats as normal.
A short window of about 7 days reacts quickly. It follows a real rise in fitness within a week, and it forgets a sick week just as fast. It also adapts to strain you would rather notice. Three hard weeks in a row slowly drag a 7-day baseline down, and each new low night starts to look ordinary. A long window of 30 to 60 days holds steady through a hard block and shows a sustained slide as a run of low scores. The cost is lag. After a lasting change, such as a big fitness gain or a move to altitude, a 60-night median takes weeks to catch up. It keeps reporting that you are above or below normal because your normal has moved.
Research practice combines the two. A short average, 3 to 7 days, stands for your current state. A longer reference, several weeks, stands for your normal. The Plews and Buchheit studies compare a 7-day Ln rMSSD average with a band from a longer baseline. The intervention studies above used 10 days to 4 weeks of reference data before making decisions. Whatever the window, it needs enough measured days. Plews's 3-days-a-week minimum is for a weekly average, and a median over a long window needs a good fraction of its days filled or it falls back on a handful of nights.
04Changing devices or metrics mid-window
A baseline assumes every value in it was measured the same way. Switching devices breaks that assumption, even between two good devices. Miller and colleagues (2022) compared six wearables with ECG over one night in the sleep lab in 53 adults. Average RMSSD bias ranged from −4.5 ms for WHOOP to −22.4 ms for Garmin. Most devices showed proportional bias, overestimating low HRV and underestimating high HRV, so the gap between two devices depends on where your own HRV sits.
Switching metrics is worse. Apple Health stores SDNN, while WHOOP and Oura report RMSSD. The two measure different parts of heart rhythm, and the ratio between them shifts with age, heart rate and recording length, so no fixed factor converts one to the other.
Here is what a median-and-MAD baseline does when you switch mid-window. As an illustration with made-up numbers, suppose your old device read about 45 ms, the new one reads about 33 ms for the same physiology, and your baseline window is 60 nights. On the first nights with the new device, a normal night sits almost 3 spread units below baseline and scores as a collapse. As new nights accumulate, the mixed window has two clusters, so the MAD roughly doubles and every z-score shrinks toward zero. The median crosses over only when new-device nights make up about half the window, around night 30. Scores stay distorted until the old nights have fully aged out after 60 nights. Firmware updates that change an HRV algorithm, or moving a watch to the other wrist, can cause a smaller version of the same effect.
The fix is simple. Keep one device and one metric for the life of a baseline. If you must switch, expect a transition period as long as the window, or restart the baseline from the switch date.
If you can, wear the old and new devices together for two weeks before the switch. Suppose the old device averaged 55 ms and the new one 62 ms over the same nights. You now know the offset, and the jump on the first night of the new device will not read as a change in you. The numbers are illustrative.
05When to restart a baseline
A new device is one of several events that change what normal looks like. Significant illness, travel across several time zones, a break from training of two weeks or more, and a new medication that affects heart rate all qualify. After any of them, mark the date, rely on how you feel and how sessions go, and give the new baseline time to fill. Outside any app the same logic applies. A week of data tells you little, and a month tells you your range.
A baseline is only as clean as the conditions that produce it. Wear the same device in the same position, at the same time of night, with a similar bedtime routine. If you take a morning reading, take it before caffeine and before you stand up. Changing conditions make the readings more variable, and a wider spread hides real change.
06How Titan builds baselines
Titan builds two different baselines, one for Recovery and one for Stress. Everything below comes from the current iOS code.
Recovery
Each night gets two inputs. The first is overnight HRV, the mean of the natural log of every HRV reading in your main sleep session, or from 22:00 to 10:00 when Health has no sleep session. The second is sleeping heart rate, the lowest 30-minute rolling average of heart rate in the same window. Titan keeps readings from one device per night, preferring the device that recorded your sleep. It drops HRV readings outside 3 to 300 ms and heart rate outside 25 to 220 bpm. A night with both inputs is a measured night, and only measured nights enter the baseline. A night missing HRV or sleeping heart rate never does.
Your baseline is the median of the measured nights in your Recovery window, 60 nights by default. Only nights before the one being scored count, so a new night never changes an earlier night's score. A baseline needs at least 7 measured nights covering at least 25% of the window, so 15 nights on the default window.
Titan measures distance from the median in units of your own spread:
z = (tonight − median) / max(1.4826 × MAD, floor)
HRV floor = 0.04 on the log scale, about 4 percent
Heart rate floor = 1 bpm
z is clamped to ±3The floor stops a run of near-identical nights from turning a tiny wobble into a huge score swing. The clamp stops one broken reading from pinning the score. HRV works on the log scale, so a 4 ms drop from 30 ms counts for more than the same drop from 80 ms.
Suppose your window held nights like the 15 in the earlier example. The median is 45 ms and the scaled MAD works out to about 0.10 on the log scale, roughly a 10 percent step. A 40 ms night sits 1.15 steps below baseline. With sleeping heart rate at baseline, that night scores 19. The score formula weights HRV twice as heavily as sleeping heart rate, and a night at both baselines scores 56. The Recovery score explainer walks through the full curve.
Titan stores each night's inputs separately for SDNN and RMSSD, so switching the HRV method never mixes the two scales in one baseline. Titan does not reset the baseline when you change watches or rings, so the transition effect described above applies. A score becomes final at noon on the day you wake. Titan keeps rereading that night's data for another day after the cutoff, so a watch that syncs late still reaches later baselines while the frozen score stays as it was.
Stress
Stress uses a fixed 60-day baseline. It takes the median of your daily average HRV and the median of the daily resting heart rate Apple Health records, and it needs at least 36 of the 60 days for each. Every 5-minute stretch of your day compares heart rate with the resting baseline and HRV with the HRV baseline as ratios. Stress has no spread term, so the percentage change drives the score directly. The Stress score article covers the screens, and the baselines help article covers baselines inside the app.
07Common mistakes
Reading the population chart instead of your trend. A 35 ms HRV says little on its own. Several nights at 35 ms when your median is 50 ms say a lot. The HRV entry and the resting heart rate entry cover what moves each input.
Reacting to one night. Day-to-day noise in HRV runs around 12 percent, several times the smallest change that matters. Look for a run of low nights, especially with sleeping heart rate rising at the same time.
Using a short window through a long hard block. A 7-day baseline absorbs three weeks of accumulating fatigue. If you are building toward overreaching on purpose, a longer window shows how far below normal you have gone.
Treating a higher baseline as the goal. Plews and colleagues (Sports Medicine, 2013) point out that elite athletes have gained fitness while their HRV fell, and that HRV can drop even as resting heart rate falls, a pattern called HRV saturation. The baseline tells you where you are relative to your recent normal. It does not tell you which direction is better for your training phase.
Mixing devices or metrics. Wear one device to bed, keep one HRV method, and give the baseline a full window to settle after any change.
08The practical takeaway
Your baseline answers one question, whether tonight is unusual for you. Build it from a median and a spread that can ignore a bad reading, keep the window long enough to see a slow decline, and measure it the same way every night. When a score looks wrong, check the baseline before you question your body. A new device, a change of metric or a window full of sick days explains many surprising scores.
09References
- Al Haddad H et al. (2011). Reliability of resting and postexercise heart rate measures. International Journal of Sports Medicine 32(8):598-605. https://doi.org/10.1055/s-0031-1275356
- Altini M, Plews D (2021). What is behind changes in resting heart rate and heart rate variability? A large-scale analysis of longitudinal measurements acquired in free-living. Sensors 21(23):7932. https://doi.org/10.3390/s21237932
- Buchheit M (2014). Monitoring training status with HR measures, do all roads lead to Rome? Frontiers in Physiology 5:73. https://doi.org/10.3389/fphys.2014.00073
- Fraser CG, Harris EK (1989). Generation and application of data on biological variation in clinical chemistry. Critical Reviews in Clinical Laboratory Sciences 27(5):409-437. https://doi.org/10.3109/10408368909106595
- Hopkins WG (2000). Measures of reliability in sports medicine and science. Sports Medicine 30(1):1-15. https://doi.org/10.2165/00007256-200030010-00001
- Javaloyes A et al. (2019). Training prescription guided by heart-rate variability in cycling. International Journal of Sports Physiology and Performance 14(1):23-32. https://doi.org/10.1123/ijspp.2018-0122
- Kiviniemi AM et al. (2007). Endurance training guided individually by daily heart rate variability measurements. European Journal of Applied Physiology 101(6):743-751. https://doi.org/10.1007/s00421-007-0552-2
- Leys C et al. (2013). Detecting outliers, do not use standard deviation around the mean, use absolute deviation around the median. Journal of Experimental Social Psychology 49(4):764-766. https://doi.org/10.1016/j.jesp.2013.03.013
- Miller DJ et al. (2022). A validation of six wearable devices for estimating sleep, heart rate and heart rate variability in healthy adults. Sensors 22(16):6317. https://doi.org/10.3390/s22166317
- Plews DJ et al. (2012). Heart rate variability in elite triathletes, is variation in variability the key to effective training? A case comparison. European Journal of Applied Physiology 112(11):3729-3741. https://doi.org/10.1007/s00421-012-2354-4
- Plews DJ et al. (2013). Training adaptation and heart rate variability in elite endurance athletes, opening the door to effective monitoring. Sports Medicine 43(9):773-781. https://doi.org/10.1007/s40279-013-0071-8
- Plews DJ et al. (2013). Evaluating training adaptation with heart-rate measures, a methodological comparison. International Journal of Sports Physiology and Performance 8(6):688-691. https://doi.org/10.1123/ijspp.8.6.688
- Plews DJ et al. (2014). Monitoring training with heart rate-variability, how much compliance is needed for valid assessment? International Journal of Sports Physiology and Performance 9(5):783-790. https://doi.org/10.1123/ijspp.2013-0455
- Vesterinen V et al. (2016). Individual endurance training prescription with heart rate variability. Medicine and Science in Sports and Exercise 48(7):1347-1354. https://doi.org/10.1249/MSS.0000000000000910
- Wu A et al. (2025). Biological variation of corrected QT and QRS electrocardiogram intervals, interpreting results of drug-induced prolongation. Western Journal of Emergency Medicine 26(4):978-983. https://doi.org/10.5811/westjem.33602
