5,000of 5,000 rows in scope
Statistics

Statistical analysis

A two-proportion z-test on conversion, plus a Welch t-test on session duration. Nothing here is asserted beyond what the sample supports.

Experiment verdictSignificant

Variant B wins the split test

Variant B converted 160.5% better than Variant A. With a p-value of 0.0000, the sample provides strong evidence that this difference is real and not random noise.

Recommendation: Ship Variant B. At the observed 160.5% relative lift, rolling out the winning design should raise conversions on the same traffic volume.

Winning variant
Variant B
Conversion difference
+8.67% pts
Relative lift
+160.5%
Confidence
100.0%
p-value
< 0.0001
Sample size
5,000
p-value
< 0.0001
Two-tailed, α = 0.05
z-score
10.354
Pooled standard error
Confidence
100.00%
1 − p
Observed power
100.0%
For the observed effect
Absolute difference
+8.67% pts
B − A
Relative lift
+160.55%
Difference ÷ baseline
95% CI of difference
+7.04% → +10.30%
Unpooled standard error
Required sample / arm
180
80% power at this effect size

Confidence intervals

Wilson 95% intervals per arm — overlap means the difference is not yet resolved

Variant A5.40% [4.58% , 6.35%]
Variant B14.07% [12.75% , 15.49%]

Secondary test — session duration

Welch's t-test (unequal variances) on mean session duration

Mean session duration is 243.3s for Variant B versus 241.7s for Variant A — a difference of +1.6s (t = 0.47, p = 0.6387). This engagement gap is not statistically significant on the current sample.