AI Max's automated optimization makes traditional A/B testing harder — the algorithm learns and adjusts, which can mask the effect of a variable you're trying to test. Google's Campaign Experiments feature solves this by creating controlled split tests within your account's traffic.
How Campaign Experiments Work in AI Max
Campaign Experiments split your campaign's eligible traffic between a base campaign and an experimental campaign. Traffic allocation is controlled (e.g., 50/50 or 70/30 split) and randomized at the user level — the same user consistently sees either the base or the experiment, not both.
This randomization is critical for valid A/B testing. Without user-level consistency, you'd measure a mix of "users who saw both campaigns," which creates comparison noise.
What you can test with Campaign Experiments:
- Bidding strategy (target CPA vs. maximize conversions vs. target ROAS)
- Budget levels (different daily budgets for the same campaign)
- Asset group configurations (different headlines, images, landing pages)
- Target CPA or ROAS values (testing looser vs. tighter targets)
- Audience signal combinations
What you cannot isolate cleanly:
- Individual asset performance (AI Max mixes assets, so testing one headline requires controlling all others)
- Placement-specific performance (you can't limit an experiment to one channel)
- Time-of-day performance (the split is across all traffic simultaneously)
Setting Up an AI Max Campaign Experiment
In Google Ads: Campaigns > Experiments > Create Experiment.
- Select "A/B test" as experiment type
- Choose your base AI Max campaign
- Set traffic split (50/50 is standard for equal statistical power; 80/20 if you want to protect base performance)
- Define the experiment duration (minimum 30 days; 45-60 days for accounts with lower conversion volume)
- Make your test change to the experiment campaign
The experiment campaign is a copy of the base campaign with your modification applied. Do not make changes to the base campaign during the experiment period — this invalidates the comparison.
The Most Impactful Tests for AI Max
Test 1: Target CPA reduction
Current tCPA $80 → Experiment tCPA $65. Does lower tCPA reduce conversion volume significantly, or does AI Max adjust efficiently? Many accounts run tCPA higher than necessary because they haven't tested how far they can push it.
Test 2: Smart bidding strategy change
Base: Target CPA | Experiment: Target ROAS (if you have revenue data). For e-commerce, switching from tCPA to tROAS often improves revenue efficiency once you have sufficient conversion value data.
Test 3: URL expansion on/off
Base: URL expansion enabled | Experiment: URL expansion disabled. Some accounts see significantly higher CVR with URL expansion disabled because AI Max's expanded URLs include lower-quality pages. This test quantifies the tradeoff.
Test 4: New asset group theme
Base: Rational headlines ("Save 30% on...") | Experiment: Emotional headlines ("Stop losing leads to..."). AI Max can't tell you which theme outperforms across your account without a controlled test.
Statistical Significance in AI Max Experiments
Google automatically calculates statistical significance for Campaign Experiments and shows whether the difference between base and experiment is statistically significant (typically at 95% confidence).
For valid results:
- Minimum sample size: At least 100 conversions in EACH of the base and experiment campaigns before drawing conclusions
- Minimum test duration: 30 days (to account for weekly seasonality patterns)
- No external changes: No major bids, budgets, or creative changes in the account during the test
The June 2026 reporting deletion (https://yositeup.com/blog/google-ads-reporting-data-deleted-june-2026) created a data irregularity period. Experiments started before June 2026 that included the deletion period may have compromised comparison data. Start fresh experiments post-July 2026 if your previous experiments ran through the deletion period.
Testing After DSA Migration
The DSA to AI Max migration (https://yositeup.com/blog/google-ads-dsa-ai-max-migration-february-2027) created natural before-after comparisons, but these aren't controlled experiments. DSA performance and AI Max performance differ due to:
- Different campaign types (not just different settings)
- Different learning phase history
- Changes in the competitive landscape at migration time
After migration, use Campaign Experiments to test AI Max configurations against each other — not AI Max against DSA. The pre-migration DSA data is not a valid control for post-migration AI Max experiments.
Interpreting Experiment Results
When the experiment ends, Google shows:
- Primary metric comparison (conversions, CPA, or ROAS depending on campaign goal)
- Statistical confidence level
- Recommendation: "Apply experiment" or "Keep base campaign"
Apply the experiment if:
- Statistically significant improvement in primary metric
- Secondary metrics (CTR, impressions, conversion rate) support the conclusion
- The improvement is practically significant (not just statistically significant — a 2% improvement in CPA may not justify switching)
Keep the base if:
- No significant difference (null result is still useful — it tells you the variable doesn't matter)
- Experiment shows worse performance
Run another experiment if:
- Results are directional but not significant (you need more volume)
- External factors (seasonality, competitor activity) may have confounded the results
Advanced Testing: Factorial Experiments
For accounts with high conversion volume (500+/month), factorial experiments test multiple variables simultaneously:
- Asset group A: Rational headlines + CTA-focused landing page
- Asset group B: Rational headlines + trust-signal landing page
- Asset group C: Emotional headlines + CTA-focused landing page
- Asset group D: Emotional headlines + trust-signal landing page
This 2x2 design reveals not just which headline or landing page performs better, but whether there's an interaction effect (rational headlines might perform better with CTA pages; emotional headlines might need trust signals).
Campaign Experiments support multiple asset groups within one experiment, making factorial designs possible. The July 2026 ToS (https://yositeup.com/blog/google-ads-tos-july-2026-ai-automation-what-changed) did not restrict experimental configurations, so factorial experiments remain a valid testing method.
Experiments vs. Asset Performance Ratings
AI Max provides asset performance ratings ("Best," "Good," "Low") which are a form of internal testing. Campaign Experiments and asset performance ratings serve different purposes:
- Asset ratings: Relative performance within one asset group's mix. Fast feedback but no control group.
- Campaign Experiments: Controlled comparison between configurations with statistical rigor.
Use both: asset ratings for rapid creative iteration within an existing structure; Campaign Experiments for strategic decisions about bidding, targeting, and campaign structure.
AI Max Shopping experiments (https://yositeup.com/blog/google-ai-max-shopping-replacing-performance-max-2026) are particularly useful for testing product-level coverage decisions — which products to include in AI Max vs. standard Shopping campaigns, and at what tROAS.
Top comments (0)