Product design brief, July 2026

All twelve apps answer how am I. None of them answers what actually works for me.

Proof is an N-of-1 evidence engine. It does not report how you are. It reports which of the things you have tried actually changed your body, by how much, and how sure it is, including when the honest answer is nothing, or not yet.

Working name Proof Core object the Verdict Success metric questions closed, not sessions Built to be needed less

A teardown of twelve consumer health apps closed with six findings the whole category shares. Four of them are correct and we took them. Two of them are the reason the category is stuck, and we refused them. Both refusals are more expensive than they look.

01
The moat is the data, not the model Accept, different stream
Sensors belong to Whoop and Apple. Labs belong to Function. Food belongs to MyFitnessPal. Passive history belongs to Google. Proof collects none of them as a primary asset. It collects the label nobody is defending: what you deliberately changed, when you started, when you stopped, and whether you actually did it. The intervention log is the only write path in the app; everything else is read-only and imported.
02
Compress value into one number, then explain it Reject at the home screen
A composite answers how am I, the question we have deliberately conceded to hardware that measures it better, and the teardown records its own failure twice: score anxiety at Whoop, users gaming their own logs at Youper. A composite is also unfalsifiable, which is fatal for a product whose whole claim is falsifiability. The home screen is the open questions board and a countdown to an answer. The single number returns only where it is earned and local: inside one Verdict, in the outcome's own units, with an interval. A number you can be wrong about is worth more than a number you cannot.
03
Constrain the conversation on purpose Accept, different mechanism
Ada, Woebot and Wysa constrain by deleting the text field. Proof constrains by scope: conversation exists only inside one Verdict, it may reference only this user's own data and trials, and it may only speak in the closed verdict vocabulary. There is no chat tab, because a free text tab is the door. An out-of-scope question routes to escalation, not to a hedged answer.
04
Make model output editable, not authoritative Accept, one layer deeper
Cal AI lets you edit the model's answer. Proof lets you edit the model's question. Tap the Signal name on any Verdict and the whole trial re-runs retroactively against a different Signal, using the nights already banked, and resolves again in front of you. The evidence base is editable and the conclusion is not. Every exclusion writes to a visible audit row. A product that lets you edit the conclusion is a mood board.
05
Solve the cold start or lose the user in week two Accept, as a hard gate
Our value accrues over time harder than anything in the twelve, because a Verdict needs two windows. So the first session is retrodictive: Backfill reads the imported history for natural experiments that already happened, a training block that started and held, a bedtime that moved by 45 minutes and stayed moved, a fortnight that went missing. The first session ends with a real Verdict about the user's own past, or with a plain statement that their history is too thin and the single import that would fix it.
06
Regulation shapes the copy, not the capability Reject the framing
Accurate about the category, and a terrible rule to inherit. Every entrant calls itself wellness while shipping the same capability, and the constraint surfaces as hedging bolted on at review. Proof goes the other way: restrict the capability so the copy can be blunt. The claim class is narrow by construction, a change in a quantity you supplied over a window you defined. No disease name, no prognosis, no treatment advice, in any tier. Because the claim is that narrow, the copy carries no hedging furniture: uncertainty is an interval, which is information, not a disclaimer, which is noise.

A dashboard is a pre-AI artifact. It exists because software could aggregate but could not reason, so the only honest thing it could do was show you the data and let you draw the conclusion. Once the reasoning layer is good enough, the correct core object stops being a state and becomes a personal causal model you can interrogate.

The unit of record is not a measurement. It is an intervention-outcome pair: what you changed, and what your body did about it. Four nouns carry the whole product, and any screen that renders none of them should be cut.

The Change

A thing you started, stopped, or shifted, with a start date and an end date.

Always a behaviour or an intake you control. Magnesium glycinate 400mg at 21:00. Last coffee before 14:00. Lifting Tuesday and Thursday. Never a goal, never an intention.

A Change is stored as blocks, not as a start date, and that is the whole reason reading your history beats a simple before-and-after. A single before and after is confounded with time by construction: the season moved, your fitness moved, your motivation moved, and you knew which side you were on. Real life is full of natural crossovers, the fortnight you stopped and the week you travelled, and finding those stops is what turns sixteen months of passive data into an alternating design nobody had to be asked to run. Marta's lifting is six blocks, not one boundary. Without the stops there is no Verdict here, only a pattern.

The Signal

A measurable thing your body does, pulled passively or supplied by one tap a day.

Every Signal carries its own noise floor, the night-to-night swing measured from your own history. No claim is ever made smaller than that floor. Here are Marta's five Signals, with every effect she has measured so far drawn in units of her own swing. The grey band is one swing either side of zero. Everything inside it is weather:

Resting heart rate 54 bpm, swings 2
HRV 48 ms, swings 9
Deep sleep 62 min, swings 18. Charted, never adjudicated
Sleep onset 19 min, swings 12
Total sleep 7h 12m, swings 41

Two ticks clear the band and three do not. The lifting dropped her resting heart rate by 4 bpm against a swing of 2. Moving her coffee to 16:00 cost her 31 minutes of sleep onset against a swing of 12. Magnesium is at 14 minutes of total sleep against a swing of 41, the cold shower at 3 ms of HRV against a swing of 9, and bed by 22:30 at 2 minutes of total sleep against a swing of 41. Those three are the ones every other app in the twelve would have called a win. Deep sleep carries no tick at all: it is charted from her own data and it is on the refuse list, so nothing is ever tested against it.

The Verdict

Proof's judgement about whether one Change moved one Signal: a state, an effect size, and an honest interval.

The only object allowed to make a claim. It is allowed to say no. It is the thing you came for, and it is the thing every other app in the twelve leaves to you.

The Handoff

One page of evidence you carry to a human clinician: what you changed, what moved, over how many nights, with the uncertainty intact.

The product's exit door and its clinical safety valve in the same object. No verdict badges on it, no colour, no logo bigger than a line of type. Numbers, dates, and your own sentence at the top.

A Change is tested against a Signal, which produces a Verdict, and settled Verdicts compose into a Handoff.
Four nouns, three verbs, no fifth concept. If a feature needs a fifth noun, it is a different product.

The centrepiece

Marta Ferreira, 41, Dublin. Sixteen months of history, and one question she cannot answer.

Everything on this device is real state. The counters count. The night goes from six to seven. The handoff counter goes from zero to one. A second question appears. And on the last screen, a Verdict that could not be resolved gets resolved, because she moved to a secondary outcome she had pre-registered before the first night rather than waiting longer.

Every physiological number here comes from her own imported history. No composite score appears at any point, on any screen. There are no streaks, no badges, and nothing celebrates.

  1. 01Connect a source
  2. 02Read sixteen months
  3. 03Nine changes already made
  4. 04Call it before we show you
  5. 05First Verdict, and how you felt
  6. 06One tap, and the night count moves
  7. 07A second question, queued
  8. 08Answers, sorted by what to do
  9. 09The one we cannot call
  10. 10Ask a pre-registered secondary outcome, gate and all
  11. 11One page for a clinician
Try this first.
Run it to the end, then on the cold shower card tap Measure something steadier instead. Try sleep onset first and read the refusal; it was never pre-registered, so it can never carry a verdict. Then pick resting heart rate, which was. You call it before you see it, exactly as you did the first time, and the answer changes under your hand using nights already banked, with no additional waiting. That is interaction 7.3, and it is the sharpest thing in the product.
9:41Proof
Screen 1 of 7

Tap the screen. Five screens minimum are reachable by tapping; twelve states in all. The single most important tap is the last one, where an INCONCLUSIVE Verdict on the primary outcome re-resolves to NOTHING THERE on a pre-registered secondary, through a second prediction gate.

A Verdict sits at exactly one state at all times, and state is recomputed nightly. The evidence bar is the same for all of them: 14 valid nights on and 14 valid nights off of the primary Signal, and for any Change you can start and stop, at least two ON blocks and two OFF blocks in alternation. One before and one after is confounded with time by construction, so a single block caps out at SIGNAL ONLY and can never reach WORKING or BACKFIRED. The bar is published inside the product with a version number, so a critic argues with a stated threshold rather than with a black box.

1
Watching
The trial is running and the bar is not met. The quietest card in the app, because a live question should not shout. 11 more nights before this can mean anything
2
Working
Bar met, the effect beats your own noise floor, in the direction you wanted, across at least two ON and two OFF blocks. No confetti. Two bars and the nights it rests on. Resting heart rate down 4 bpm, across 6 blocks
3
Nothing there
Bar met, both windows full, the difference is smaller than the floor. Deliberately the same size and shape of card as WORKING, so a null reads as a result. Your sleep did the same thing either way
4
Backfired
Bar met, effect beats the floor in the direction you did not want. The only alert treatment in the product, used once. Sorts to the top regardless of anything else. This is going the wrong way. Worth stopping.
5
Inconclusive
The window ran, adherence was good, and the difference still cannot be separated from your own noise. Different from NOTHING THERE, and it gets a different screen. We ran all 21 nights and we still can't call it
6
Confounded
Illness, travel, or a second Change invalidated the window. The set-aside nights are struck through in the strip so you can see exactly which ones went. Too much changed at once to blame the magnesium
7
Signal only
A one-way change where a washout is impossible or unsafe: a house move, a new job, a long-acting medication. Proof reports the movement and refuses the causal claim. A pattern we can see, not a verdict we can defend
8
Retracted
A settled Verdict withdrawn, months later, because the growing family of tests makes the original claim unsafe to keep. Run enough comparisons and some of them come back positive by chance; the correction arrives when the arithmetic says so, not when it is convenient. We are taking one back
RETRACTED is the differentiator, and it is not a defect. No app in the twelve has ever come back and told a user it was wrong. This one does, in its own words, without hiding the earlier claim: "In March we told you cold showers lifted your HRV. You have run 19 experiments since then, and with that many comparisons, a result that size shows up by chance more often than we were accounting for. We are moving that one back to no clear effect. We would rather correct ourselves than let you keep believing something we cannot stand behind." A retracted result is removed from the clinician handoff entirely. A product that revises its own history quietly cannot be trusted about the present; a product that revises it loudly can.
Why re-asking the question is methodology, not outcome shopping. The sharpest gesture in the product looks, at first glance, like the oldest fraud in science: a trial fails on its outcome, so the analyst goes looking for one that worked. Proof closes that door at the moment a Change is marked, not at the moment it resolves. The Mark sheet declares one primary Signal and automatically pre-registers two or three secondaries, printed on the sheet before the first night is collected: Primary: sleep onset. Also pre-registered: total sleep, resting heart rate. Declaring more than one outcome is not free, and the app says so on the same screen: every declared outcome has to clear a stricter bar than a single one would. So when a question stuck at INCONCLUSIVE on HRV is re-asked of resting heart rate, the user is switching to an outcome that was written down in advance, and the resolved card carries the receipt: SECONDARY OUTCOME, PRE-REGISTERED 1 MAY. Ask for a Signal that was never declared and the app charts it and refuses the Verdict, in those words, and offers to start a new pre-registered trial instead. That refusal is the feature. It teaches, through one interface gesture and without a word of statistics, the difference between a question you can answer and a question you have already answered by choosing it.

Commitment one

No Verdict is ever revealed before you record what you think it will say.

One screen, three taps, enforced on the server so the answer cannot reach the device first. This is a gate, not a prompt, and there is no way past it.

Before we show you the answer MAGNESIUM 400MG AT 21:00, tested against total sleep 14 nights on, 14 nights off, in 4 alternating blocks. What do you think it did? [ Helped ] [ Did nothing ] [ Made it worse ] How sure are you? [ Guessing ] [ Fairly sure ] [ Certain ]

The reveal then opens with the comparison rather than the result: You said it helped. It did. Or the more valuable one, You said it helped. It did nothing. The gap between your prediction and the evidence, tracked over time, is the capability metric. Calibration is the real output; the Verdict is only how it gets measured. Without this gate there is no learning mechanism, no graduation trigger, and Proof collapses into a better-argued Oura.

Commitment two

Every resolving Verdict asks about the outcome we did not measure, and prints the disagreement.

One question, one tap, above the effect size and at equal typographic weight. The verdict screen's hierarchy is the product's entire statement about what counts as real, and users read it as such whether or not we intend them to.

And how did you actually feel? [ Better ] [ Same ] [ Worse ] Resting heart rate down 4 bpm, across 6 blocks.
Your sleep improved and you felt worse. Both are real. We measured one of them.
Rendered on the card when the body and the experience disagree. The product refuses to resolve it.

This exists because of a failure mode that scales with success rather than with failure. A product that only ever proves things about sleep, HRV, glucose, weight and resting heart rate teaches an implicit lesson over three years: the things worth believing are the things with a sensor attached. The user keeps the rule that is proven and drops the Thursday dinner that the rule inconveniences, because nothing in their life produces a verdict saying the dinner mattered. Every individual choice is rational. The aggregate is a life optimised toward the sensor-legible, and the losses are invisible by construction. So the product says it out loud, once at the start and once at graduation: we can only prove things about what we can measure. That is a limit of our instruments, not a statement about what matters in your life. Most of what matters will never show up here.

Rule one outranks every other line in the design and every commercial argument that will be made against it: Proof is decision support for a person and their clinician, and it must never behave as though it could be a substitute for one.

The refuse list

Signals too noisy or too poorly measured to carry a Verdict. Proof will chart every one of these. It will never adjudicate them, however clean the arithmetic looks:

  • Sleep stage minutes
  • Single-timepoint blood biomarkers
  • Spot blood pressure
  • Single-exposure glucose response
  • VO2max estimates
  • Bioimpedance body fat
  • Metabolic age composites
  • Cognitive tests with a practice effect
  • Anything across a firmware change
  • Vendor readiness and recovery scores

The deep sleep verdict is the most seductive and least defensible claim in the category. Consumer stage classification disagrees with polysomnography, and the error is not random with respect to the things people change.

Escalation is free forever, for everyone

Red flags, emergency numbers and the clinician handoff are free on every tier, including expired and cancelled accounts, including users who never paid, and never behind account creation. The teardown documents a competitor putting human escalation behind an upgrade; that is the single worst decision in the set and we do not repeat it. This binds the business model, not just the interface: it is a stated cost of doing business, never a tier feature. If Proof cannot make money without charging distressed people for the exit, Proof should not exist.

Escalation copy is static, versioned, human-written and human-translated. No model is ever permitted to rewrite an emergency string. Every helpline number is verified against the operator's own published information before launch and re-verified on a schedule, because a wrong number on that screen is the worst bug this product could ship.

The verdict floor

Hard-coded refusals at the output layer, not model instructions, because model instructions fail and hard-coded refusals do not. No verdict about a disease, its presence, absence, progression or risk. No verdict about a medication, a pharmacologically active supplement, or a dose. No verdict that recommends starting, stopping, changing or timing any treatment. No prognostic claim: no biological age, no life expectancy, no years added. No verdict in pregnancy, on children, or on fertility. No comparative verdict against other users.

When a user logs a medication change as their intervention, which will happen constantly, the answer is precise rather than a shrug: Proof accepts the log and does not run the experiment. It marks the date on every chart, excludes and extends every other running trial, and writes a dated line into the clinician handoff. If the signals improve after someone stops a medicine, the app says nothing about the medicine. That silence is the correct clinical behaviour.

The clinician handoff, and the test it has to pass

One side of A4, black on white, and it has to earn its place in four seconds with a GP who has been handed patient printouts before and learned that they waste the seven minutes. Every cell carries an n and a date range. Every row states what was measured, what was estimated, and what was self-reported. The exclusions are printed, so a clinician who can see what was thrown away can judge what remains. No disease name, no differential, no recommendation, and no sentence containing suggests, consistent with, or may indicate. Those are clinical reasoning verbs and they do not belong to us.

Before launch the page goes cold in front of at least twenty practising GPs, and it is iterated until a clear majority say it saved them time. That is written into the plan as a launch gate. If they say it wastes their time, the feature is wrong and the copy will not save it.

The residual objection, quoted, from the clinician who signed the rest of it

"False reassurance is the deepest risk in the product, and it is structural. A person receives NOTHING on a change they made, or simply sees no red flag, and concludes they are fine. They delay a consultation. We can write the absence of a flag is not a clean bill of health on every screen and it will not survive contact with how people actually read. The whole escalation architecture is tuned to detect a small set of patterns in a handful of signals; the overwhelming majority of serious disease produces nothing this app can see. A person who trusts Proof more than their own sense that something is wrong is worse off than a person with no app. I do not have a fix. I have wording, and wording is weak medicine."

That objection stands unresolved and it is printed here rather than answered, because a safety section that only reassures is marketing. Two more sit beside it: a product built on measurement and personal optimisation is close to ideal apparatus for a person with disordered eating, and the mitigations are partial; and the text classifier that catches the most dangerous inputs sits behind a text field that people in crisis do not use.

The named asset is a corpus of pre-registered intervention-outcome pairs, adherence-weighted, attached to individual physiological baselines. Not a model. Not an interface. A labelled dataset of deliberate human acts and their measured consequences, which does not exist at consumer scale because nobody has built the cheap collection loop for it.

Two properties make it an asset rather than a log. Every pair is pre-registered: the outcome measure and the duration are fixed before any data arrives, which is what separates evidence from retrospective storytelling. And every pair carries an adherence weight, because a trial abandoned on day four is a different object from one that was completed, and the corpus is only as good as its willingness to say so.

Within one user it compounds through calibration, not volume. Every completed trial, including every null, measures how variable that person's outcome is week to week, and that variability sets how many observations the next trial needs. After a handful of trials, Proof can say this question needs eighteen nights rather than a generic six weeks. Time to verdict falls monotonically. The user's switching cost is not their data, which is portable and which we export; it is their calibrated trial economics. Leaving means going back to generic trial lengths.

Across users it compounds through the priors library, matched on baseline variability, age band, chronotype and the intervention itself, and it produces three things the user sees: how many observations this question usually needs, a starting estimate that shortens a first trial, and the base rate of nothing, the share of matched users for whom this intervention did nothing. No other product in the category has any incentive to compute that last number, and it is the one users most want.

Why can't Google ship this next quarter

What Google can copy, quickly
  1. The interface. In a sprint.
  2. The statistics. N-of-1 trial design is published, taught, and unremarkable. There is no algorithmic secret here.
  3. The passive data. They already have more of it than we ever will, plus Fitbit-lineage history and platform-level context.
  4. Distribution. They can put a button in front of an installed base we cannot reach by any other route. If the only thing between us were engineering, we lose.
What they structurally cannot or will not
  1. The output is anti-engagement. The most common Verdict retires a behaviour and reduces sessions. A team inside a platform can build this once; it cannot survive the second planning cycle in which it is shown to reduce its parent surface's usage.
  2. First-party causal claims are a liability surface they avoid. Moving from "sleep trending down" to "this thing you did caused this change in your body" is a step change for a company that size. We absorb it by narrowing the claim class. Their risk is not proportional to the claim; it is proportional to Google.
  3. The label requires volunteered disclosure to a trusted recipient. The corpus is built from people telling us what they are taking, quitting, or trying, including things they are not proud of. Health data plus an advertising company is a trust problem no interface designs its way out of. They can build the feature and still not get the labels.
  4. Whoop and Oura face the sharper version. Their product is the composite number, and a product that tests whether the number responds to anything can prove the number does not matter. They are the least likely fast followers, not the most.

The half-life, stated plainly. The concept has a half-life of roughly one product cycle. Once this gets meaningful attention, assume a credible clone is announced within twelve to eighteen months; the idea is not protectable and the methodology is public. The corpus is the only durable part, and its half-life is a function of one number: pre-registered trials completed per active user per quarter. If that number is healthy, a follower needs both the users and the calendar time, because trials take weeks and cannot be bought. If it stalls, the moat evaporates and this is a nice interface with a thin database. The single metric this company should be run on is completed trials, not installs.

This product is built to be needed less, and that sentence is worthless unless a number moves when it happens. So the headline number is one the company wins by making go down.

Unaided Call Accuracythe primary metric

The share of your own pre-reveal calls that the evidence subsequently agreed with, over decisive Verdicts only, rolling over the last ten in a domain. Exact match, no partial credit, because partial credit is where metrics go to be gamed.

Engagement measures whether you came back. Retention measures whether you paid. This measures whether the model moved from the machine into the person, and it is indifferent to whether you opened the app today.

Trials to Competencethe company north star

The median number of decisive trials a user completes in a domain before their call accuracy crosses and holds the competence threshold. Lower is better. It is the only headline number proposed for a consumer health company where the company wins by making the number fall, and it is expensive to fake.

Three counter-metricsnever reported apart

Base-rate excess. Call accuracy minus the accuracy of always guessing that user's own most frequent state. Users can learn that nothing is the common answer; that is not learning their body.

Trial difficulty index. If accuracy rises while the questions get easier, the number is being gamed and the correct reading is that the product got worse.

Reassurance open rate. Sessions where the user opened the app, gave no input, got no new Verdict, and looked at an old one. That is the signature of checking rather than learning. It should fall as tenure rises. If it rises, every other number is lying to you.

Graduationa screen, not a value

Per domain, never global. It fires when eight of the last ten calls in that domain were correct, base-rate excess is positive, at least one of those correct calls was a null or a backfire, three rules are held, and ninety days have passed since the first trial in that domain, so graduation can never become an onboarding trophy.

Three choices at equal weight, no default highlighted: step back to monthly, keep it as it is, close the file. Nothing congratulates. Nothing offers to share an achievement. Nothing upsells a second domain in the same breath. The user's rules leave with them as a one-page plain document, readable with no app, no account and no subscription, designed to be printed and stuck inside a cupboard door.

The standardasked every six months

Do the people who have been with us longest need us least, and can they prove it by calling their own results correctly before we show them? Both clauses are load-bearing. If session frequency rises with tenure, the answer is no regardless of what anyone believes about the mission. And a user who opens the app less because they lost interest is not a graduate; call accuracy is what separates the two.

The pricing has to be shaped for this rather than fighting it. Monthly billing punishes success, because the user cancels in the exact month the product worked and the company quietly learns to delay verdicts. Annual billing decouples the payment moment from the usage moment and removes the incentive to slow a trial down. Beside it sits a dormant tier at a token price that keeps a graduated user's baselines and calibrated trial economics warm, so returning costs nothing but a tap. Design for return, not for retention.

  • What it is

    A design specification and a working prototype, not a company and not a clinical safety case. Nothing here has been tested on a real user. The thresholds are reasoned defaults, the escalation copy has never been read by a frightened person, and the evidence standard has never met real device data.

  • Resolved during review, recorded

    The first build tested magnesium against deep sleep, which the clinical refuse list bars from ever carrying a Verdict. The seed data and the refuse list were written by different seats and the conflict was real. It was resolved by moving that trial onto total sleep, a signal whose noise floor this product can actually measure. Deep sleep still appears under Body as raw data, because hiding a person's own data is paternalism; what it may never do is carry a claim. The episode is left on the page because a product about evidentiary honesty should show its own corrections.

  • Unresolved, two

    The most likely way this dies is quiet: individual effects of most consumer interventions are small relative to individual day-to-day noise, so run honest statistics and the majority of trials return nothing or not yet. The integrity feature and the unpleasantness are the same feature. The early signal to watch is whether users start a second trial after their first null.

  • Unresolved, three

    The open trial with a countdown is the strongest engagement hook in the concept and it arrived dressed as integrity. Every gradient in a subscription business points toward longer trials and tighter evidence bars, and it will arrive as a series of individually defensible decisions about rigour. The guardrail is that median trial length and the ratio of not-yet to decisive verdicts are treated as regressions when they rise.

  • What would have to be true

    That a health-literate adult will tell a piece of software what they are actually taking; that enough trials resolve to something for the product to feel like an answer machine rather than a disappointment machine; that a GP handed the one page says it saved them time; and that a physician will put their name on the trigger rule set. The first is the bet. The other three are testable before launch, and each of them is written into the plan as a gate.