How long should an A/B test run? There is no universal number of days that works for every experiment.
The right duration depends on your required sample size, baseline conversion rate, minimum detectable effect (MDE), statistical significance, statistical power, and daily eligible traffic.
A simple way to estimate test duration is:
Test Duration = Required Sample Size ÷ Average Daily Eligible Traffic
However, reaching the required sample does not always mean you should immediately stop the test. Your experiment should also capture normal variations in user behavior, including weekday and weekend patterns.
What determines A/B test duration?
Four statistical inputs have the biggest impact on the required sample size:
Baseline conversion rate: Your current conversion rate before launching the test.
Minimum detectable effect: The smallest improvement worth detecting.
Statistical significance: The evidence threshold used to reduce the risk of false positives.
Statistical power: The probability of detecting a real effect when one exists.
A common testing setup uses a 5% significance level, equivalent to 95% confidence, and 80% statistical power.
The smaller the improvement you want to detect, the more traffic you generally need. A test designed to identify a small 1% improvement can require substantially more traffic than one designed to identify a 20% improvement.
Why sample size matters
Sample size should be calculated before launching your experiment.
If you test with too little traffic, you may fail to detect a genuine improvement. This creates an underpowered test.
On the other hand, running a test much longer than necessary can waste traffic and delay implementation of a potential winner.
Your sample size calculation should therefore happen during test planning, not after the experiment has already started.
Don't stop an A/B test as soon as it becomes significant
One of the most common A/B testing mistakes is stopping an experiment immediately after the platform reports a significant result.
Test results fluctuate as new visitors enter the experiment. An early result can look impressive simply because the initial sample does not represent your normal audience.
This practice, often called peeking, can increase the chance of false-positive conclusions.
Instead, define your sample size and expected test duration before launch. Continue monitoring the experiment for technical problems or severe performance drops, but avoid using temporary statistical fluctuations as a reason to declare a winner.
Why running for a full week matters
Even if your experiment reaches its calculated sample size quickly, you should consider whether it has captured normal traffic patterns.
For example, customer behavior can differ between:
Weekdays and weekends
Working hours and evenings
Promotional and non-promotional periods
New and returning visitors
A very short experiment may capture an unusual traffic pattern rather than typical customer behavior.
For many experiments, running through at least one to two complete weeks provides a better representation of normal behavior. Longer durations may be necessary when the business has strong seasonality, long conversion cycles, or lower traffic.
What if the test is inconclusive?
An inconclusive A/B test does not automatically mean the variation failed.
It can mean:
The actual effect was smaller than the MDE.
The experiment did not have enough statistical power.
The metric has high natural variability.
The hypothesis was incorrect.
The audience or traffic conditions changed during the experiment.
Review the confidence interval, sample size, baseline conversion rate, and MDE before deciding what to do next.
If the potential impact is still commercially important, you may need more traffic or a redesigned experiment.
A practical A/B testing process
A reliable testing process can follow these steps:
Define the primary conversion goal.
Establish the baseline conversion rate.
Choose a realistic MDE.
Set your significance level and statistical power.
Calculate the required sample size.
Estimate duration using your daily eligible traffic.
Build and QA the experiment.
Run the test without stopping early because of temporary results.
Analyze the final results.
Implement the winner or use the findings to develop the next hypothesis.
The key takeaway
There is no universal answer to “How long should an A/B test run?”
Your experiment should run long enough to collect the required sample size and capture representative user behavior. Calculate the required sample before launch, estimate the duration from your eligible traffic, and avoid stopping simply because the results look promising early.
A disciplined testing process helps teams make decisions based on reliable evidence instead of short-term fluctuations.
Read the complete guide:
https://www.brillmark.com/how-long-should-you-run-an-a-b-test-sample-size-significance-duration/
BrillMark helps businesses plan, develop, QA, and launch statistically sound A/B tests. Learn more about A/B test development, CRO support, and experimentation services at BrillMark.
Top comments (0)