Small experiment tool

A/B Test Delta Calculator

Enter the users and conversions from each group. The calculator focuses on the difference between Treatment and Control—not on whether two separate confidence intervals happen to overlap. Inputs stay in this browser.

Four inputs

Enter your result

A / B

Use unique users as the sample unit. Conversion counts cannot be larger than user counts.

Control · A

The version users would have seen before.

Treatment · B

The version you are testing.

The p-value uses a two-sided pooled two-proportion z-test. The 95% delta interval is a Newcombe interval built from Wilson score intervals. The calculator calls evidence clear only when p < 0.05 and the interval excludes zero. It reports direction only unless each arm has at least five expected successes and failures under the pooled null.

Delta firstCalculating

Treatment is ahead

Update the inputs to see how clear the difference is.

Control rate
Treatment rate
Absolute lift
Relative lift
P-value
95% delta CI
Conversion rate
Control
Treatment
Look at the delta, not the overlap.

The pooled p-value and Wilson-based Treatment minus Control interval answer related questions. This tool reports both and requires them to agree before calling the evidence clear.

Term guide

Plain English
Control · A
The original version, used as the baseline for comparison.
Treatment · B
The new version or change being tested against Control.
User / sample
One unit in the experiment. Here, it means one unique user.
Conversion
The outcome you are counting, such as a signup or purchase.
Conversion rate
Conversions divided by users. For example, 100 / 1,000 = 10%.
Delta
Treatment rate minus Control rate. A positive delta means B is higher.
Absolute lift
The delta expressed in percentage points, such as +2 pp.
Relative lift
The delta divided by Control. A move from 10% to 12% is a 20% relative lift.
Confidence interval · CI
A range showing plausible values for an estimate, given this sample.
95% delta CI
A Wilson-based uncertainty range for Treatment minus Control, shown in percentage points. It does not collapse at rates of 0% or 100%.
P-value
How surprising this difference would be if the two versions truly had the same rate. Smaller values mean more evidence against that assumption.
Statistically significant
The pooled z-test uses p < 0.05. Because that test and the Wilson-based interval are different approximations, this tool calls evidence clear only when p < 0.05 and the interval excludes 0, with adequate expected counts.
Two-sided test
A test that checks for either an increase or a decrease, rather than only looking for a win.
Two-proportion z-test
The approximate statistical method used here to compare two conversion rates.
Sample size
The number of users included in a group. Total sample size is Control users plus Treatment users.
5% level
The p-value threshold used by this quick calculator. The separate interval is shown at 95%; this tool requires both methods to agree.

Results are calculated locally in your browser. No experiment data is sent anywhere.