Foundational · Jun 9, 2026 · 6 min read

How accurate is a phone at this, really?

A phone resting on a nightstand at night, its microphone facing a sleeping person — the only sensor SomniSense uses to estimate breathing irregularities.

It's a fair thing to be skeptical about. A clinical sleep study straps a dozen sensors to you — airflow, oxygen, chest movement, brain waves. SomniSense uses one microphone, sitting on the nightstand, two feet away. So when someone asks me "how accurate can that possibly be," I don't think they're being rude. I think they're asking the right question.

The honest answer is: more accurate than you'd guess, less accurate than a sleep lab, and — this is the part that matters — you can't describe it with a single number. Anyone who gives you one number is hiding something.

Why "98% accurate" is a marketing sentence, not a measurement

Here's the trick almost every sleep app plays. They quote one big round number — "98% accurate!" — and they don't tell you accurate at what. Accurate at noticing a sound happened? Easy, anything can do that. Accurate at telling a snore apart from a passing truck? Harder. Accurate at counting the right number of breathing pauses across a whole night? Harder still, and a completely different question.

The two numbers that actually matter pull in opposite directions:

  • Sensitivity — of the real events that happened, how many did we catch? Miss too many and we'd undercount your night.
  • Precision — when we flag an event, how often was it really one? Flag too eagerly and we'd cry wolf, inflating your number with false alarms.

You can make either one look great by sacrificing the other. An app that flags everything has perfect sensitivity and useless precision. The only honest disclosure is to publish both, and let you see the tension. A single blended "accuracy" number is what you report when you'd rather people didn't look closely.

What our numbers actually are

So, plainly. For 1-second snore-event detection, the five-seed means are 91.67% sensitivity and 89.01% precision. For 200-second breathing-event windows, the compressed model reports 88.49% accuracy and 88.06% F1. Those are research-cohort classification benchmarks, not a personal accuracy guarantee for tonight's SRI or BRI.

I'm deliberately not going to unpack every one of those here, because that's a longer, more careful read, and I'd rather you have it in full than in a blog-sized summary. The number-by-number version lives on the accuracy page, and the study design and the model behind it — how a phone gets to those numbers at all — is on the research page.

The honest part: where it's weaker

The events we miss are mostly the borderline ones — the short, shallow pauses sitting right on the threshold. We miss them on purpose, because we tuned the model to avoid false alarms. The practical consequence: if your real number is a 12, we might tell you 10. But if we tell you 18, you can trust it's at least 18. I'd rather under-promise than inflate.

It's also weaker in a noisy room, a partner who snores louder than you, a fan blowing straight at the phone. And it was validated on adults — not kids. None of that is in the marketing copy of most apps. It should be.

If you'd rather see your own night than read about someone else's —

It's free on both stores. Put the phone on the nightstand tonight; the report is there when you wake up.

First 7 days of Pro are free · then $7.99/mo or $49.99/yr · cancel before day 8 and you pay nothing.