Earlycall
Method

The maths, and the data it was calibrated on

Every number this product quotes should survive the question says who? This page is the answer: the statistics the engine runs on, how its guarantee is tested, and every dataset behind the figures — with citations, so you can check them without asking us.

The statistics

A standard significance test assumes you look once, at a sample size fixed before you start. Its 5% false-positive rate is a promise about that single look. Teams do not work that way — dashboards get checked daily, and every extra look is another chance for noise to cross the line. Checked weekly, the practical false-alarm rate rises to roughly 25–50% rather than 5% — the exact figure depends on how many looks the test lives through and when they land, which is precisely the problem: the guarantee you were promised no longer describes the procedure you ran.

Earlycall uses an always-valid confidence sequence — a time-uniform, empirical-Bernstein bound. Where a fixed-horizon interval is valid at one pre-planned moment, a confidence sequence is valid at every moment simultaneously. Stopping early, stopping late, or checking hourly cannot inflate the error rate, because the guarantee was never conditioned on when you looked. The practical consequence is the whole product: the first moment the interval clears your bar is a moment you can act on.

The empirical-Bernstein form matters for speed. It adapts to the variance actually observed rather than assuming the worst case, which is what buys the early call instead of merely a safe one.

How the guarantee is tested

A claim about false positives is only worth what its adversarial test is worth. The engine ships with a calibration battery of A/A experiments — both arms identical, so every "winner" is by construction a false alarm — decided under continuous peeking, the exact condition that breaks the standard method.

0 / 300

False ships across 300 A/A runs of 200,000 units each, decided under continuous peeking against a 5% guarantee. Every run graded A; none required the integrity gate to block a ship.

Regenerable: make certify reruns this tier with hard assertions, so an engine change that breaks a quoted claim breaks the build. The battery also covers continuous metrics, three-way segment interactions, adversarial delivery (30% redelivery, 20% outcome-before-exposure), and a sample-ratio-mismatch gate that must refuse to ship a real 20% lift on corrupted assignment.

The production data

Priors, effect bands and replay evidence come from public experiment archives rather than from intuition. Each is cited; each is downloadable by anyone who wants to re-run the analysis.

DatasetScaleWhat it calibrates
Upworthy Research Archive
Matias, Munger & Wright, Scientific Data 8:195 (2021)
32,487 A/B tests Copy and creative effect bands — the widest spread measured. Also the source of three provable winners in 7,644.
ZOZO Open Bandit Dataset
Saito et al. (2020), CC BY 4.0
13.7M impressions The hour-by-hour replay on the home page, and personalization effect bands.
Criteo uplift
Ad-exposure experiments
13.9M users Marketing and exposure effect bands.
ASOS digital experiments 78 experiments Incremental UX bands at a mature product — why polish rarely moves conversion.
REES46 e-commerce events 2 months of sessions Traffic seasonality and the testability calendar: which questions open a window and when.
Hillstrom email 64,000 customers Metric-choice and covariate pricing: what CUPED is actually worth on real data.

What is open, and what is not

Being straight about this, since the rest of the page asks you to trust numbers:

Open: every dataset above is public and cited. The findings are reproducible by anyone with the data — the analysis scripts are plain Python and Go, and the calibration battery is regenerable on demand. The method is standard published statistics, not a proprietary trick.

Not open: the engine source is closed. So the honest status of the A/A figures is regenerable by us on every build, not independently auditable by you today. If that matters for your decision, ask and we will share the raw battery output and the analysis scripts behind the corpus tables — you can re-run the dataset findings yourself against the public archives above, which is where the effect claims actually come from.

Where the claims appear

Every verdict the product gives cites its own basis: direction bands carry the corpus and its size, the certified risk floor names the regime it was measured in, and a prior that has not met its fitness bar is labelled safety-only rather than quietly used. If you find a number in this product that cannot be traced back to this page, that is a bug — tell us.