Most failed A/B tests fail before they launch: the hypothesis was vague, the metric was chosen afterwards, or the test was stopped the moment the dashboard turned green. A one-page brief, completed and agreed before anything is built, prevents all three.
Copy the template below into your ticket system or experiment log.
Experiment brief
ID: EXP-<number>
Owner: [name]
Status: Draft / Approved / Running / Concluded
1. Observation
What did you see, and where did you see it? Link the evidence.
Example: Session recordings and form analytics show that a large share of mobile visitors who open the quote form abandon it at the "company size" field.
Evidence sources can include analytics funnels, heatmaps, session recordings, on-site surveys, support tickets and usability tests. An opinion in a meeting is not an observation.
2. Hypothesis
Use this structure:
Because we observed [observation],
we believe that [change]
for [audience / segment]
will cause [expected behaviour change].
We will know this when [primary metric] changes in [direction].
3. Metrics
| Type | Metric | Definition |
|---|---|---|
| Primary (one only) | e.g. quote form completion rate | Completed quote forms / sessions that viewed the form |
| Guardrail | e.g. lead quality | Share of leads accepted by sales |
| Guardrail | e.g. page performance | No regression in LCP or INP on the tested template |
| Secondary (diagnostic) | e.g. field-level drop-off | Not used to call the result |
Decide the primary metric before launch. Reporting whichever metric happened to move afterwards is how false wins get shipped.
4. Design
- Variants: Control vs.
- Audience and targeting: device, market, language, traffic source
- Traffic split: usually 50/50
- Unit of randomisation: user (preferred) or session
5. Sample size and duration
- Estimate the required sample with a calculator, using your current baseline conversion rate and the smallest effect worth detecting for the business.
- Run for whole weeks (at least one, usually two or more full business cycles) to avoid day-of-week bias.
- If the traffic needed exceeds what the page receives in a reasonable time, do not run the test. Use qualitative research or make the change based on evidence instead.
6. Stopping rule
Write it down: "We will stop when each variant reaches [n] visitors and at least [x] full weeks have passed." Do not stop early because a result looks significant on day three - repeated peeking inflates false positives unless your tool uses a method designed for it.
7. QA checklist before launch
- [ ] Variant renders correctly on mobile and desktop, in every language the page supports
- [ ] No flicker of the original content before the variant loads
- [ ] Analytics events fire identically in both variants
- [ ] The test does not break forms, checkout or consent banners
- [ ] Exclusion rules work (internal traffic, bots)
8. Result and decision
| Field | Entry |
|---|---|
| Outcome | Win / Loss / Inconclusive |
| Primary metric result | [effect and confidence interval] |
| Guardrails | [any regressions] |
| Decision | Ship / Iterate / Discard |
| What we learned | [one or two sentences, regardless of outcome] |
Keeping a learning log
Store every concluded brief - including losses and inconclusive tests - in one searchable place. Over time the log becomes the most valuable CRO asset you own: it stops the team re-running the same idea and shows which kinds of changes actually move your audience.
Further reading
- For a broader view of how a structured, funnel-wide conversion programme is run, see how Gorilla Vibe runs conversion programmes.
Top comments (0)