Earlycall
Free · no signup The call sheet

Your test didn’t settle it. Make the call anyway.

Most tests never reach significance, and most teams ship on a hunch and write it up as a win. Paste four numbers and get the honest version: what each way of being wrong actually costs you, how often a coin flip looks exactly like your result, and a decision you can defend on Thursday.

Try it on:
Use my own numbers
A Your current version
B The challenger
Your bar
No signup. Works with the numbers from any tool — Optimizely, VWO, GA4, a spreadsheet — and every peek keeps the 5% guarantee.
Free tool The past-wins audit

Were the wins you already shipped real?

Three provable winners in 7,644 public tests. Point the same method at yours. Nothing is stored.

One row per finished test, any column order: name, control visitors, control conversions, treatment visitors, treatment conversions. Try it on a sample file →

Free tool The metric advisor

Which question can your traffic answer?

Most wasted experiments were unanswerable before they started. Here’s which of yours can finish.

Free tool The gut-call check

Can’t prove it? Decide anyway — with the risk priced.

Most ideas can’t be proven at most teams’ traffic. Decide deliberately instead of pretending — odds and a priced downside, no traffic spent.

The evidence Replayed hour by hour

When could you have called it?

One production experiment, 13.7 million impressions. Both methods watching the same week.

Show the working — hour-by-hour statistics
Earlycall’s guaranteed range Old calculator’s 95% range Observed lift

Real production data: ZOZO Open Bandit Dataset (CC BY 4.0, Saito et al. 2020), replayed through the same engine that answers the calculator above. Traditional side: the standard two-proportion z-test on identical inputs. The A–F ribbon under the chart is Earlycall’s evidence grade — its live answer to “can I trust the data itself right now?”

The evidence Replays of public archives

Replayed at scale

The wolf-cries

Run 20 tests where both versions are identical — every “winner” is a false alarm. Watched weekly like a real dashboard, a significance calculator’s false-alarm rate reaches 25–50% depending on the looks, not the 5% it promised.

The archive
1/7,644

Of 7,644 winners declared across the largest public archive of real A/B tests, exactly one was provably better than its runner-up. Read the analysis →

Straight answers A/B testing questions

Questions people actually ask

Can I check an A/B test before it finishes?

With Earlycall, yes. It uses always-valid confidence sequences, whose 5% false-positive guarantee holds at every look rather than at one pre-planned sample size. With a standard significance calculator, no: it assumes you check exactly once, and checking repeatedly is what inflates its error rate.

Why does peeking break statistical significance?

A p-value below 0.05 means that if there were truly no difference, a result at least this extreme would show up less than 5% of the time — in a single, pre-planned test. Each additional look is another chance for noise to cross the line. Checked weekly the way real dashboards are watched, the practical false-positive rate is far above 5% — simulations of typical peeking behaviour put it anywhere from roughly 25% to 50%, depending on how often and how long you look.

How long should I run an A/B test?

Long enough for the smallest lift worth shipping to be distinguishable from noise at your traffic — which for rare conversions can be longer than the test is worth. Earlycall's metric advisor computes this before you spend traffic, and will tell you when the honest answer is that your traffic can never settle the question.

What is always-valid (anytime-valid) inference?

A family of methods whose error guarantee holds at every moment you look, so stopping early or late does not invalidate the result. Earlycall uses a time-uniform empirical-Bernstein confidence sequence. The practical consequence is that the first moment the interval clears your bar is a moment you can act on.

Is a p-value below 0.05 enough to ship?

Only if you looked exactly once, at a sample size you fixed in advance, and nothing about the data collection was corrupted. In a replay of 32,487 real A/B tests from the Upworthy Research Archive, of 7,644 declared winners exactly three were provably better than their runner-up under an always-valid bound \u2014 and two of those three differed only in the photograph.

What if my traffic is too low to run A/B tests?

Then most of your decisions are judgment calls, and the useful thing is to price them rather than dress them as tests. Earlycall's gut-call check gives the odds a change like yours clears your bar (from public corpora of real experiments) and the typical cost of being wrong, without spending any traffic.