What Is a Baseline and Why Does It Take Two Weeks to Build?
Date Published
Aug 18, 2026Time to Read
7 minA baseline isn't a single average. It's your own range for one measure: a center point, plus a width that describes how far your readings normally sit from that center. Two weeks is roughly what it takes to pin down both parts. The studies that ask this question directly land on thresholds of five to fourteen usable days, and calendar time has to leave room for the days you miss.
Key Takeaways
- A baseline has two parts, a center and a width, and it is the width that decides whether today's reading counts as a deviation or as ordinary scatter.
- The reliability research converges on a working floor rather than a magic number, and two weeks of calendar tracking is what reliably delivers the number of usable days those studies ask for.
- Measures that describe your variability settle more slowly than measures that describe your average, so parts of a baseline keep sharpening past the two week mark.
Why an average alone can't be a baseline
Most people hear the word and picture an average. Actually, an average alone can't do the job a baseline is for.
Say your nightly heart rate variability has averaged 52 ms over the past month, and last night came in at 44 ms. Does that mean anything? Not yet. The answer depends on something the average doesn't contain: how much your nights normally differ from each other. If your readings usually land within a few milliseconds of 52, last night is a departure. If they routinely swing across a 20 ms span, last night is Tuesday.
Same number, two different meanings. That's the part a single average hides.
So a baseline is a description of your usual state under repeatable conditions, and it needs two numbers: one for where your readings center, one for how widely they spread. Clinical laboratories formalized this long before wearables existed. Biological variation data, meaning the within-person and between-person scatter of a measurement, are what get used to calculate reference change values and to build personalized reference intervals for an individual patient (Sandberg et al., 2023). A reference change value answers exactly the question above: how far does a repeat result have to move before the move is real?
That's also the mechanical reason a baseline takes days rather than minutes. A center point converges fast. A width has to wait for your readings to show how far they actually travel.
How long does it take to establish a baseline?
This is one of the few questions in wearable research where the studies agree with each other.
In 107,144 nights recorded by 1,041 working adults aged 21–40 using a consumer tracker, 3 nights were needed for a good weekly estimate of mean total sleep time and 5 nights for a very good one, at intraclass correlations of 0.60 and 0.80 (Lau et al., 2022). Across roughly 2 million nocturnal HRV readings from more than 21,000 wearable users, at least five of seven nights were required before a weekly index of day-to-day HRV fluctuation agreed acceptably with the full week value (Grosicki et al., 2026). Different measures, different cohorts, similar floor.
Five usable nights isn't five calendar nights, though. And short observation windows flatter themselves. When 79 participants wore a Fitbit Flex daily for about a year, a resampling approach showed that the conventional practice of extrapolating from brief test-retest windows can substantially overestimate real long-term reliability, and that 6, 8 and 10 valid days were actually needed to reach 0.80 reliability for light activity, sedentary time and moderate to vigorous activity (Hilden et al., 2023). Their recommendation was 7 to 10 valid days. In lived practice that means more than 7 to 10 calendar days.
Missing nights are why. If you've ever left the watch on the charger overnight, gone to bed without it after a late one, or come back from a trip to find the week looks thin, you've already made the kind of gap this is about. Below that five-of-seven threshold the weekly estimate stops agreeing with the full week, so a gap in the series does more than thin your data. It moves the reference your later readings get compared against. Fourteen days of tracking with a few gaps is what reliably delivers ten usable ones, and that arithmetic is most of the reason HaloScape asks for a two week window before it starts reading anything into a single day.
Where two weeks stops being enough
Two weeks builds a usable range. It doesn't finish one. This is the part that rarely makes it into the pitch, and the evidence is specific about which parts lag.
In 336 real-world continuous glucose monitor users with type 1 diabetes tracked over 90 days, 14 days of sampling was sufficient for mean glucose and time in the 70–180 mg/dL range, while the coefficient of variation required 28 days and time below 54 mg/dL required 35 (Akturk et al., 2025). The sleep dataset above shows the same ordering: the 5 nights that gave a very good weekly mean grew to 11 and 18 nights once the target was monthly sleep duration variability rather than monthly sleep duration itself.
Averages settle first. Spread settles second. Rare low-end events settle last. So the honest version of the two week claim is this: two weeks is the floor at which a personal range becomes worth comparing against, not the point at which it's finished. The number describing how much you normally bounce inside that range keeps sharpening for another two to three weeks. Why a personal reference beats a population figure in the first place is a separate argument we make here.
A baseline is the boring part. Two weeks in which nothing appears to happen, which is exactly the stretch everyone wants to fast-forward through to reach the insight. But the insight is made out of the boring part. A deviation is a claim about a range, and with no range, a reading has nothing to deviate from. What you get for the waiting isn't a score. It's the ability to look at an ordinary night and know that it's ordinary, and to recognize the night that isn't.
References
- Sandberg S, Carobene A, Bartlett B, Coskun A, Fernandez-Calle P, Jonker N, Díaz-Garzón J, Aarsand AK. Biological variation: recent development and future challenges. Clinical Chemistry and Laboratory Medicine. 2023;61(5):741–750. PMID: 36537071. doi:10.1515/cclm-2022-1255
- Lau T, Ong JL, Ng BKL, Chan LF, Koek D, Tan CS, Müller-Riemenschneider F, Cheong K, Massar SAA, Chee MWL. Minimum number of nights for reliable estimation of habitual sleep using a consumer sleep tracker. Sleep Advances. 2022;3(1):zpac026. PMID: 37193398. doi:10.1093/sleepadvances/zpac026
- Grosicki GJ, Carter JR, Laursen PB, Plews DJ, Altini M, Galpin AJ, Fielding F, Hippel WV, Chapman C, Jasinski SR, Beattie UK, Holmes KE. Heart rate variability coefficient of variation during sleep as a digital biomarker that reflects behavior and varies by age and sex. American Journal of Physiology: Heart and Circulatory Physiology. 2026;330(1):H187–H199. PMID: 41309064. doi:10.1152/ajpheart.00738.2025
- Hilden P, Schwartz JE, Pascual C, Diaz KM, Goldsmith J. How many days are needed? Measurement reliability of wearable device data to assess physical activity. PLoS ONE. 2023;18(2):e0282162. PMID: 36827427. doi:10.1371/journal.pone.0282162
- Akturk HK, Sakamoto C, Vigers T, Shah VN, Pyle L. Minimum Sampling Duration for Continuous Glucose Monitoring Metrics to Achieve Representative Glycemic Outcomes in Suboptimal Continuous Glucose Monitor Use. Journal of Diabetes Science and Technology. 2025;19(2):345–351. PMID: 37747124. doi:10.1177/19322968231200901
Give a new baseline two weeks of same-window, same-device days before you read anything into a single reading, then keep recording, because the part of the range that tells you how much movement is normal for you is still getting sharper in week three and week four.