By Brooke Quinn, MSc (Oxon), RPSGT, CSSC
Every night in sleep labs across the globe, technicians sit in front of monitors displaying a complex, messy reality that looks nothing like the sleek, color-coded graphs delivered by popular wellness apps. The in-lab screen is a cacophony of physiological data: raw brain waves (EEG), rapid eye movements (EOG), chin muscle tone (EMG), airflow, breathing effort, oxygen saturation, and heart rhythm.
We review this data in 30-second segments, or "epochs," assigning each a label: wake, N1, N2, N3 (deep sleep), or REM. While standardized rules (AASM criteria) provide a framework, human sleep is rarely so binary. Signals overlap, transitions are fluid, and even two highly trained sleep technologists may interpret the same ambiguous segment differently.
Yet, for millions of consumers, the nuance of sleep science has been condensed into a single, alluring number: an "accuracy percentage." This reductionist approach is now at the center of a burgeoning legal and ethical conflict, highlighted by a recent class-action lawsuit against Oura. As we navigate the era of the "quantified self," it is time to unpack why these percentages often obscure more than they reveal.
Main Facts: The Anatomy of a Legal and Technical Dispute
The recent class-action lawsuit filed against Oura (Surber v. Oura, Inc.) alleges that the company misled consumers regarding the efficacy of its sleep-stage tracking capabilities. At the heart of the complaint is the discrepancy between marketing claims—specifically the assertion of "95% sleep staging accuracy"—and the reality of how these devices function when compared to the "gold standard" of clinical polysomnography (PSG).
The core of the issue is the definition of "accuracy." In the clinical world, accuracy is a multidimensional metric requiring sensitivity, specificity, and a deep understanding of epoch-by-epoch agreement. In the marketing world, it has become a shorthand for "how often the device gets it right." The conflict arises because "getting it right" is an inherently subjective goal depending on whether one is measuring total sleep time, sleep onset latency, or the intricate architecture of sleep stages.
Chronology: From Lab Bench to Living Room
The rise of consumer sleep technology has followed a rapid, often chaotic trajectory:
- Pre-2015: Sleep tracking was largely the domain of high-end clinical equipment or rudimentary actigraphy devices that could only track movement.
- 2015–2020: The proliferation of photoplethysmography (PPG) sensors in wrist-worn and finger-worn devices allowed for the monitoring of heart rate variability (HRV) and peripheral capillary oxygen saturation (SpO2). Companies began applying machine-learning algorithms to these signals to "predict" sleep stages.
- 2023–2024: Validation studies began to appear in peer-reviewed journals, showing that while consumer devices were becoming increasingly sophisticated, their performance varied wildly depending on the population tested—healthy sleepers vs. those with chronic insomnia or sleep apnea.
- August 2026: A formal class-action complaint is filed against Oura, alleging deceptive trade practices. The lawsuit ignited a firestorm of debate, pitting tech enthusiasts who value longitudinal trends against clinicians who fear the normalization of unvalidated diagnostic data.
Supporting Data: Understanding the Inference Gap
To understand why the "95% accuracy" claim is controversial, one must distinguish between measurement and inference.
An Oura ring does not record brain waves. It records pulse-wave information, heart-rate patterns, skin temperature, and movement. It then uses a proprietary algorithm to infer the likely state of the brain. This is not inherently fraudulent—physiological changes (like a drop in heart rate and body temperature) are indeed correlated with deep sleep. However, inference is not invention.
The data provided by Oura itself reflects this complexity. While they report roughly 90% to 96% agreement for distinguishing sleep from wakefulness (a binary task), that number drops significantly—to between 76% and 79%—when the device is tasked with the more complex, four-stage classification (Wake, Light, Deep, REM).

A 2025 study of sleep-lab patients further highlights the volatility of these metrics. In this clinical setting, the device achieved 85% accuracy for sleep-versus-wake, but only 53% accuracy across the four distinct sleep stages. More strikingly, the study found that while the average "total sleep time" error was low (an overestimate of about 12 minutes), the individual variance was massive—ranging from a 54-minute underestimate to a 97.5-minute overestimate. For a patient relying on this data to manage a medical condition, that 97-minute error is not a "minor discrepancy"—it is a fundamental misrepresentation of their sleep health.
Official Responses and Industry Standards
In its response to the legal challenges, Oura has emphasized the distinction between their consumer-facing "wellness" focus and clinical diagnostic tools. They maintain that their algorithms are continuously updated and that their marketing materials reflect the high performance of their sensors in the context of healthy, non-clinical populations.
However, the American Academy of Sleep Medicine (AASM) has remained consistent in its position: consumer sleep technology may enhance the patient-clinician interaction but cannot, and should not, replace validated diagnostic instruments like polysomnography. The AASM warns that these devices lack the regulatory rigor of medical-grade hardware, particularly regarding how they handle irregular sleep patterns or sleep disorders.
Implications: The Clinical and Personal Fallout
The implications of this debate extend far beyond the courtroom. We are witnessing the emergence of "orthosomnia"—a term coined by clinicians to describe patients who develop an unhealthy obsession with achieving "perfect" sleep scores. When a patient sees a low deep-sleep score on their app, it can trigger anxiety, which in turn causes further sleep fragmentation.
1. The Erosion of Trust
When consumers are promised clinical-grade accuracy and later discover that their device failed to detect an hour of wakefulness, the trust in both the technology and the healthcare system is damaged.
2. Clinical Overload
Clinicians are now regularly presented with printouts of wearable data. The challenge is not to dismiss the data, but to translate it. A clinician must determine: Is this patient reporting daytime sleepiness because of the data they see, or because of a physiological issue? Does the longitudinal trend of "decreased REM" correlate with a new medication, or is it a sensor error?
3. The Need for Transparency
The industry must pivot toward "transparent performance." A meaningful accuracy statement should not be a single percentage. It should be a disclosure of:
- The Population: Who was the device tested on? (Healthy adults vs. patients with OSA).
- The Task: Is the accuracy for sleep/wake or for stage classification?
- The Limitations: What is the margin of error for specific sleep stages?
Conclusion: Toward a More Nuanced Future
Useful does not have to mean certain. The true value of a wearable lies in its ability to track longitudinal patterns—the "big picture" of sleep regularity, timing, and duration. These metrics can be invaluable for identifying the impact of travel, shift work, or lifestyle changes.
The legal battle against Oura is, in many ways, a symptom of a larger, systemic growing pain in the wearable tech industry. As these devices move from "fun gadgets" to "health tools," the marketing language must evolve from absolute claims of "accuracy" to a more honest disclosure of "utility."
We do not need wearable companies to be sleep labs. We do need them to be transparent about the boundaries of their algorithms. Trust is not earned by hiding the messiness of biology behind a perfect percentage; it is earned by inviting the user into the reality of what we can—and cannot—know about the dreaming mind.
References
- Surber v. Oura, Inc., No. 3:26-cv-08686 (ND Cal filed 20 Aug 2026).
- Robbins R, et al. "Accuracy of three commercial wearable devices for sleep tracking in healthy adults." Sensors (Basel). 2024.
- Svensson T, et al. "Validity and reliability of the Oura Ring Generation 3… compared to multi-night ambulatory polysomnography." Sleep Med. 2024.
- Oura Team. "Standing behind our science: how Oura measures sleep and validates accuracy." 2026.
- Herberger S, et al. "Performance of wearable finger ring trackers for diagnostic sleep measurement in the clinical context." Sci Rep. 2025.
- Khosla S, et al. "Consumer sleep technology: An American Academy of Sleep Medicine position statement." J Clin Sleep Med. 2018.
- Baron KG, et al. "Orthosomnia: Are some patients taking the quantified self too far?" J Clin Sleep Med. 2017.
