Before launching a test, know how many visitors each variant needs to make the result reliable. Enter current CVR, minimum detectable effect and daily traffic to get required sample and estimated test duration.
An A/B test without pre-calculated sample size is a coin flip. You can see a 15% difference between variants and be looking at pure noise; you can see a 2% difference and have a statistically valid winner. The only way to separate signal from noise is to start with a pre-calculated sample.
Sample size depends on three decisions: how strict you want to be with false positives (confidence level), how sure you want to be of detecting a real effect (statistical power), and how small the effect you want to detect (MDE).
Formula (per variant)n = 2 × p̄(1−p̄) × ((Zα/2 + Zβ) ÷ Δ)²
Where p̄ is the expected average CVR across control and variant, Δ is the expected absolute difference, Zα/2 = 1.96 for 95% confidence and Zβ = 0.84 for 80% power. Intimidating on paper; the calculator does it for you.
The starting point. The CVR of your control today. Lower baseline → larger sample required to detect differences.
The smallest lift you want to reliably detect. 20% MDE on a 3% base means detecting a move to 3.6% or higher. Small MDEs are ambitious: they demand huge samples.
Probability of avoiding a false positive. 95% (α = 0.05) is the marketing standard. Lowering to 90% cuts sample but raises risk. Raising to 99% multiplies sample by 1.7.
Probability of detecting an effect that actually exists. 80% (β = 0.20) is the acceptable minimum. At 70% you miss 3 out of 10 real winners due to lack of sensitivity.
| Baseline CVR | MDE 10% | MDE 20% | MDE 30% |
|---|---|---|---|
| 1% | ~118,000 | ~29,500 | ~13,100 |
| 2% | ~58,400 | ~14,600 | ~6,500 |
| 3% | ~38,600 | ~9,650 | ~4,300 |
| 5% | ~22,800 | ~5,700 | ~2,550 |
| 10% | ~10,800 | ~2,700 | ~1,200 |
Approximate values at 95% confidence, 80% power, two variants.
The most expensive CRO mistake. Stopping before the pre-calculated sample raises the false-positive rate to 30-40%. Wait. Trust the math.
Each extra variant adds sample. With limited traffic, prioritise simple A/B (2 variants) over multivariate. Learnings arrive sooner.
Even if the sample is reached in 3 days, run the test at least 7 to capture day-of-week variability. Saturday users behave differently than Tuesday ones.
If your page doesn't produce statistical traffic, it doesn't mean you can't optimise: it means you should prioritise big value-prop changes, not detail tweaks. Test entire landing restructurings instead of button colors. In parallel, invest in SEO and content to grow the traffic base to the threshold where quantitative CRO becomes viable.
To know how many visitors per variant you need before the result is statistically valid. Without it, you risk stopping the test early and acting on results that are actually random noise.
Depends on baseline and effect size. With 3% CVR chasing a 20% lift you need ~9,500 sessions per variant for 95% confidence and 80% power. Smaller baseline or smaller effect → sample grows fast.
The smallest lift you want to reliably detect. A 20% MDE on a 3% baseline means detecting a move to 3.6% (or more) with confidence. Smaller MDE, larger sample required.
95% confidence (alpha 0.05): the chance of calling a variant a winner when it's not is just 5%. 80% power (beta 0.20): if the variant truly wins, you have 80% chance of detecting it. Industry standard.
Yes, with caveats. Low traffic only lets you detect big changes (high MDE), which limits what makes sense to test. Practical rule: if your page doesn't get at least 10,000 sessions/month, prioritise big value-prop changes over detail optimisations.
No. Stopping before the pre-calculated sample raises the false-positive rate (the 'peeking problem'). Run the test to the target sample, even if a winner looks clear.
At least one or two full weeks to capture day-of-week variability, even if the sample is reached earlier. Never less than 7 days. A 3-day test can be biased by different weekend behaviour or one-off campaigns.
We design end-to-end SEO programs for brands that want to lower their paid dependency, scale predictable organic traffic and dominate Google plus the new AI search surfaces.
SEO strategy, technical SEO, content, authority and AI visibility (GEO) under one senior team. SEO tied to pipeline and revenue, not vanity metrics.
Squad embedded in Slack, Notion and Linear. Weekly sprints, a 12–24 month roadmap and executive reporting that connects every organic move with CAC, leads and revenue.
Compounding organic growth, lower paid dependency, growing share of voice in LLMs and an acquisition channel that keeps working when you stop paying for clicks.
Technical, content, authority and LLM visibility audit. Benchmark vs. competitors and opportunity quantified in traffic and revenue.
Topic universe prioritised by intent, 12–24 month traffic projection and a measurement model wired to business outcomes.
Technical fixes, briefs, content, on-page, authority and GEO shipped every week. Embedded operation, not deliverables that sit in a PDF.
Organic KPIs tied to pipeline and CAC. Monthly iteration on what is actually moving the business, not on what climbs in Search Console.
We will send you a free SEO diagnostic with the real organic opportunities for your domain and a projection of how much you could cut paid spend while keeping the same lead volume.