This page is updated every two months with current best practices for Google Ads experiments, drafts and A/B testing. Testing is how you separate real improvements from noise, but a poorly designed experiment can send you confidently in the wrong direction. Each update draws on our own experience plus authoritative industry sources and verified real-time research. Bookmark this page and check back for the latest experiment best practices. Each update includes worked examples with the arithmetic shown.
Last updated: 6 August 2026
In This Guide
- Executive Summary
- Benchmarks & Numbers at a Glance
- Why Test
- Experiments, Drafts & A/B Tests
- Designing a Valid Experiment
- Sample Size & Significance
- What to Test
- Reading Results
- Applying or Rolling Back
- Testing with Automation
- Common Mistakes to Avoid
- What Changed Recently
- References
1. Executive Summary
Google Ads experiments remain the most defensible method for attributing performance change to a single cause in a paid search programme. Five principles govern effective testing in 2026.
- One variable, one metric. Every experiment must isolate a single change and declare a primary success metric before the test starts. Cherry-picking metrics after results arrive is the leading cause of false positives.[1][8]
- Volume before velocity. A campaign that cannot deliver at least 30–50 conversions per variant per month will rarely reach a trustworthy conclusion within a practical timeframe. Prioritise experiments on your highest-traffic campaigns.[5][7]
- Respect the ramp-up and conversion lag. Discard the first 7 days of data from any experiment and extend the window to cover your full conversion lag before reading results. Stopping early is the second leading cause of false positives.[8][12]
- 95% confidence is the safe default. An 80% threshold moves faster but produces more noise. Reserve the lower bar only for low-stakes, easily reversible tests where speed matters more than certainty.[1][3]
- Document every test. Record the hypothesis, split, dates, primary metric, result, and next action. Institutional memory from sequential tests compounds into durable performance gains; ad-hoc testing does not.[1][5][9]
2. Benchmarks and Numbers at a Glance
| Metric | Typical range or threshold | Applies when | Source |
|---|---|---|---|
| Recommended experiment duration — standard ad copy test | 14–30 days | Campaign has sufficient conversion volume; vendor claim | [58] |
| Recommended experiment duration — general minimum | 2–4 weeks | Most Search and Shopping campaign types; practitioner guidance | [5][3] |
| Recommended experiment duration — bidding strategy change | 60 days minimum | Switching Smart Bidding strategy; practitioner guidance | [1] |
| Recommended experiment duration — major campaign structure change | 90 days minimum | Structural overhauls such as campaign consolidation; practitioner guidance | [1] |
| Ramp-up period to discard from results | ~7 days (approximately 1 week) | All experiment types; vendor claim (Google) | [8][11] |
| Learning-phase buffer before evaluating bidding experiments | 14 days | After any Smart Bidding strategy switch; practitioner guidance | [1] |
| Algorithmic learning phase for bid-strategy switching | 7–14 days | When switching between automated bid strategies; practitioner guidance | [64] |
| Ideal conversions per variant — ad copy test | 200 conversions per variation | Ad copy A/B tests requiring high confidence; practitioner guidance | [58] |
| Minimum conversions per variant — Performance Max asset test (high-volume) | 30–50 conversions per variant | PMax asset experiments with strong campaign volume; practitioner guidance | [4] |
| Minimum conversions per variant — Performance Max asset test (lower-volume) | 50–100 conversions per variant | PMax asset experiments with lower base conversion volume; practitioner guidance | [4] |
| Minimum monthly conversions for any experiment to be viable | 30–50 conversions per month | Any campaign type before starting an experiment; practitioner guidance | [5] |
| Statistical confidence threshold — conservative default | 95% | Strategically important or expensive-to-reverse tests; practitioner guidance | [1][3] |
| Statistical confidence threshold — faster iteration | 80% | Low-stakes, easily reversible tests only; higher false-positive risk; practitioner guidance | [1][2] |
| Traffic split — balanced test | 50/50 | When volume is sufficient and maximum statistical power is needed; practitioner guidance | [3][5] |
| Traffic split — risk-reduction test | 70/30 (control/experiment) | Higher-risk changes where protecting base performance matters; practitioner guidance | [2][3][5] |
| Maximum experiment duration (initial end-date window) | 2–12 weeks from start | Standard Google Ads experiment setup; vendor claim (Google) | [60] |
| Maximum extension from current date | 12 weeks | Extending an active experiment; vendor claim (Google) | [60] |
| AI Max experiment traffic split | 50/50 (fixed, set by Google) | AI Max Search experiments; vendor claim (Google) | [4] |
3. Why Test in Google Ads
Without controlled experiments, performance changes in a Google Ads account are almost impossible to attribute reliably. Seasonality, auction pressure, Quality Score fluctuations, and Smart Bidding’s own learning cycles all move metrics simultaneously. An experiment that isolates a single variable creates a counterfactual — what would have happened without the change — that no amount of retrospective analysis can replicate.[8][13]
The business case for structured testing is compounding. Each experiment produces a documented learning that narrows the hypothesis space for the next test. Over a 12-month programme of sequential 4–6 week experiments, an account can run 6–8 well-powered tests, systematically working through bidding strategy, landing-page message match, ad copy themes, and audience signals.[1][3]
Google’s own experiment infrastructure reinforces this discipline. The Experiments page — the successor to the older drafts-and-experiments workflow — provides a single place to create, manage, and read results from all test types, with built-in significance reporting that removes the need for manual statistical calculations in most cases.[11][15]
The cost of not testing is equally concrete. Accounts that make frequent, uncontrolled changes to Smart Bidding targets, budgets, and creative simultaneously are resetting machine learning without ever knowing which change caused which outcome. This creates a feedback loop of reactive optimisation that rarely compounds into durable gains.[16][3]
Worked example
Quantifying the value of a sequential testing programme
- Setup: A Melbourne e-commerce account spending $18,000 per month on Search campaigns, generating 120 conversions per month at a blended CPA of $150 AUD.
- Numbers: At 4 weeks per test with a 1-week ramp-up discarded[8], the account can complete 10 sequential experiments in 50 weeks (10 × 5 usable weeks). Each test targets one variable. Even if only 3 of 10 tests produce a statistically significant winner at 95% confidence[1], and each winner reduces CPA by 8% ($150 × 0.92 = $138 after test 1; $138 × 0.92 = $127 after test 2; $127 × 0.92 = $117 after test 3), the compound CPA reduction is 22% over the year. At 120 conversions per month and a starting CPA of $150, that is a saving of $150 − $117 = $33 per conversion × 1,440 annual conversions = $47,520 AUD per year in recovered budget or redeployed spend.
- Decision: Commit to a formal sequential testing programme: 10 experiments scheduled across 52 weeks starting 4 August 2026, each capped at 5 weeks (1 week ramp-up + 4 weeks evaluation), one variable per test, 50/50 split, 95% confidence threshold.
- Why: Even a low win rate of 3 in 10 produces material compounding savings when each winner is applied before the next test begins.[1][3]
4. Experiments, Drafts and A/B Tests Explained
Google Ads offers several distinct testing mechanisms. Understanding what each does — and when to use it — prevents misapplication and wasted traffic.[11][15]
Drafts
A draft is a staging copy of an existing campaign. You make proposed changes to the draft without those changes going live. The draft can then be applied directly to the campaign (replacing the original) or converted into an experiment (running alongside the original). Drafts are the correct preparation step when you want to test a campaign-level change such as a restructured ad group, a new bidding strategy, or a revised keyword list.[11][15]
Campaign Experiments
A campaign experiment runs the draft variant against the original campaign simultaneously, splitting traffic according to a percentage you set. Google’s system tracks metrics separately for the control (original) and the treatment (experiment variant), enabling a statistically grounded comparison. The Experiments page is the modern entry point for creating and managing these tests.[11][15] Only one experiment can run per campaign at a time.[5]
AI Max Experiments
Introduced and expanded in 2026, AI Max experiments are a one-click test for Search campaigns that evaluates AI-powered features including search term matching and asset optimisation. Unlike traditional campaign experiments, AI Max experiments automatically split the existing campaign 50/50 without requiring you to create a campaign copy first. This is intended to accelerate the test setup and produce results faster.[4]
Performance Max Asset Experiments
Performance Max has a dedicated experiment format for testing asset groups. This workflow differs from Search campaign experiments and requires a minimum conversion threshold of 30–50 conversions per variant for high-volume accounts, or 50–100 for lower-volume accounts, before results are reliable.[4]
Broad Match Experiments
Search campaigns support a Broad Match experiment type that now provides incremental query insights, showing which additional search queries the broad match variant is capturing compared to the control. This is a meaningful reporting improvement for accounts evaluating a shift to broader match types.[3] (Note: Research 4 identifies this as a recent addition; prior documentation did not include incremental query reporting at this level of granularity.)
| Experiment type | Campaign type | Traffic split | Requires a draft? | Key use case |
|---|---|---|---|---|
| Campaign experiment (custom) | Search, Shopping | Configurable (e.g., 50/50, 70/30) | Yes | Bidding strategy, structure, keyword, audience changes |
| AI Max experiment | Search | 50/50 (fixed by Google) | No | Testing AI search term matching and asset optimisation |
| Broad Match experiment | Search | Configurable | Yes | Evaluating incremental query volume from Broad Match |
| Performance Max asset experiment | Performance Max | Configurable | No (asset-level) | Testing asset group creative variants |
| Demand Gen experiment | Demand Gen | Configurable; split immutable after setup | Varies | Creative and audience testing in Demand Gen campaigns |
Worked example
Choosing between a campaign experiment and an AI Max experiment for a Search campaign
- Setup: A Brisbane legal services account spending $9,000 per month on a single Search campaign, running Maximise Conversions bidding, generating 45 conversions per month at $200 AUD CPA. The account manager wants to test AI-powered search term matching but also wants to test switching to tCPA at $200.
- Numbers: 45 conversions per month exceeds the 30–50 per month viability threshold[5]. Two simultaneous experiments on the same campaign are prohibited — only one experiment per campaign at a time[5]. The AI Max experiment uses a fixed 50/50 split set by Google[4]; the tCPA bidding test requires a custom campaign experiment with a configurable split. Running both simultaneously is not possible.
- Decision: Run the tCPA bidding experiment first using a custom campaign experiment at 50/50 for 60 days starting 10 August 2026[1]. Schedule the AI Max experiment to begin after the bidding experiment concludes on 9 October 2026.
- Why: Only one experiment can run per campaign at a time[5], and the bidding foundation must be stable before testing AI features, which is the recommended testing order.[1][10]
5. Designing a Valid Experiment
Experiment validity depends on decisions made before the test starts. Decisions made after the test starts — changing the split, editing the campaigns, adding a new conversion action — all contaminate the result and render the data untrustworthy.[8][10]
Write a testable hypothesis
A valid hypothesis names the variable, the expected direction of change, and the mechanism. “Switching from Maximise Conversions to tCPA at $180 will reduce CPA without reducing weekly conversion volume, because the tighter target will force the algorithm to prioritise higher-intent queries” is testable. “Testing a new bidding strategy to improve performance” is not.[1][2][5]
Isolate one variable
Change exactly one thing between the control and the treatment. If you change both the bidding strategy and the headlines in the same experiment, any movement in the primary metric is unattributable. Google’s own guidance and all major 2026 practitioner sources are unambiguous on this point.[8][1][5][9]
Set the traffic split before launch and do not change it
Google’s FAQ now explicitly states that for all experiment types — including custom experiments and Demand Gen experiments — the traffic split percentage cannot be changed after setup. If you want a different split, you must create a new experiment.[2] Use 50/50 when volume is sufficient and you want maximum statistical power. Use 70/30 (control/experiment) when the change carries execution risk and protecting the base campaign’s performance takes priority.[2][3][5]
Freeze everything else
During the experiment, do not edit keywords, bids, negatives, budgets, ad copy, landing pages, or audience lists in either the control or the treatment campaign unless those changes are the thing being tested. Mid-test edits are the single most common source of contaminated results in practitioner accounts.[8][10][14]
Set the start date 3–7 days in the future
Google’s API documentation recommends setting experiment start dates 3–7 days in the future to allow time for review and approval processes.[5] This also gives you a clean cut-off date for change history review before the experiment begins.
Choose one primary metric and one guardrail metric
The primary metric decides the winner. Common choices are CPA, ROAS, or conversion value per cost. The guardrail metric catches regressions elsewhere in the funnel — for example, if conversion volume falls significantly even while CPA improves, the guardrail signals a problem. Do not add guardrail metrics after results arrive.[2][3][5]
Worked example
Writing a valid experiment hypothesis and setup document
- Setup: A Perth HVAC installation account spending $12,000 per month on Search, running Maximise Conversions, averaging 60 conversions per month at $200 AUD CPA. The team wants to test moving to tCPA.
- Numbers: 60 conversions per month is above the 30–50 per month viability threshold[5]. A 60-day bidding strategy experiment[1] at 50/50 gives each arm approximately 60 conversions per 30-day period, or ~120 per arm over the full 60 days — above the 30–50 minimum[5]. Start date: 17 August 2026 (set 7 days in advance[5]). Ramp-up period discarded: 17–24 August 2026 (7 days[8]). Evaluation window: 25 August – 16 October 2026. Primary metric: CPA. Guardrail metric: weekly conversion volume must not fall below 12 conversions per week (80% of the 15-per-week baseline). tCPA target set at $200 (matching current CPA baseline).
- Decision: Document and lock the following before launch — hypothesis: “Setting tCPA at $200 will hold CPA at or below $200 AUD while maintaining at least 12 conversions per week”; variable: bidding strategy only; split: 50/50; start: 17 August 2026; end: 16 October 2026; primary metric: CPA; guardrail: ≥12 conversions/week. No changes to keywords, ads, or landing pages during the experiment.
- Why: Naming the variable, direction, mechanism, and guardrail before launch prevents post-hoc metric selection, which is the leading cause of false positives.[1][3]
Worked example
Applying the 70/30 split for a higher-risk creative overhaul
- Setup: A Sydney financial planning account spending $22,000 per month on Search, generating 80 conversions per month. The team wants to test a complete rewrite of all RSA headlines from generic service descriptions to pain-point-led copy — a significant creative risk.
- Numbers: A 70/30 split means the control receives 70% of traffic (~56 conversions per month) and the treatment receives 30% (~24 conversions per month). At 30 days, the treatment arm has ~24 conversions — just at the lower bound of the 30–50 minimum[5]. Extending to 42 days (6 weeks) gives the treatment arm ~34 conversions, comfortably above 30. Set start date 7 days in advance: 10 August 2026. Evaluation window begins 17 August 2026, ends 20 September 2026 (42-day test minus 7-day ramp-up = 35 evaluation days). Primary metric: conversion rate. Confidence target: 95%[1].
- Decision: Set split to 70/30 at experiment creation (immutable after setup[2]); run for 42 days from 10 August 2026; do not change to 50/50 mid-test — create a new experiment if a rebalance is needed.
- Why: The 70/30 split protects 70% of monthly revenue-generating traffic while still delivering enough treatment conversions for a valid result after 42 days.[2][3][5]
6. Sample Size, Duration and Significance
The most common experiment failure in Google Ads is not a tracking error or a strategic mistake — it is ending the test too early. Early stoppage dramatically inflates the false-positive rate because conversion data from the final days of a conversion window has not yet been attributed. Google’s guidance explicitly recommends excluding the initial ramp-up period of approximately 7 days and the most recent days where conversion reporting is incomplete when judging results.[8][11]
Duration by test type
Duration requirements vary significantly by what is being tested. The table in Section 2 provides the full set of thresholds. As a working rule: ad copy tests need a minimum of 14 days and ideally 21–30 days[58]; bidding strategy tests need a minimum of 60 days[1]; major structural changes need a minimum of 90 days[1]; and all tests must cover at least one full weekly cycle to avoid day-of-week bias.[2][5]
Conversions per variant
Google’s platform does not publish a universal minimum sample-size formula, but practitioner guidance converges on the following thresholds: 200 conversions per variation is the ideal threshold for ad copy tests[58]; 30–50 conversions per variant is sufficient for high-volume PMax asset tests[4]; and a campaign generating fewer than 30–50 conversions per month is unlikely to produce a trustworthy result within a practical timeframe regardless of test type.[5]
Statistical significance
Google’s built-in experiment reporting provides significance indicators. Practitioner guidance recommends waiting for 95% confidence before declaring a winner for any strategically important or expensive-to-reverse test.[1][3] An 80% threshold is faster but materially more prone to false positives — it is only appropriate for low-stakes, easily reversible changes.[1][2] Where the two thresholds conflict in a given test, the conservative recommendation is 95%.[1][3]
Accounting for conversion lag
Conversion lag is the time between a click and the conversion event being recorded. For e-commerce with same-session purchase, lag may be under 24 hours. For B2B SaaS with a multi-week sales cycle, lag may be 30–90 days. Any experiment evaluated before the conversion lag has elapsed will undercount conversions in the most recent portion of the test window, systematically biasing results against the treatment arm if traffic ramped up during the test.[12][5] The practical fix is to extend the evaluation window by at least one full conversion lag period beyond the experiment end date before reading final results.
Worked example
Calculating minimum duration for a B2B SaaS bidding experiment
- Setup: A Sydney B2B SaaS account spending $15,000 per month on Search, generating 35 conversions per month (form fills attributed to a 14-day conversion window), average CPA $428 AUD. The team wants to test switching from Maximise Conversions to tCPA at $430.
- Numbers: Minimum duration for a bidding strategy test: 60 days[1]. Ramp-up to discard: 7 days[8]. Learning-phase buffer: 14 days[1]. Conversion lag: 14 days. Total window required: 60 days test + 14 days conversion lag buffer = 74 days. At 35 conversions per month and a 50/50 split, each arm receives approximately 17–18 conversions per month, or ~42 conversions per arm over 74 days — above the 30 minimum[5]. Start date: 1 September 2026. End date for experiment traffic: 31 October 2026. Earliest valid read on results: 14 November 2026 (allowing 14 days for conversion lag to clear).
- Decision: Set experiment end date to 31 October 2026 but do not evaluate results until 14 November 2026. Primary metric: CPA. Confidence target: 95%[1].
- Why: Reading results before 14 November 2026 would exclude conversions still in the 14-day attribution window, understating true conversion volume and inflating apparent CPA.[12][5]
7. What to Test: Bidding, Copy, Landing Pages, Audiences
The 2026 best-practice testing order is: establish a stable bidding foundation first, then improve message match through landing pages, then test ad copy themes, then layer in audience signals. This sequence matters because bidding stability is a prerequisite for clean copy and landing-page tests — an unstable Smart Bidding strategy will mask copy-lift signals with algorithmic noise.[1][4][10]
Bidding
Test bidding strategy changes in sequence, not simultaneously. For new or low-data campaigns, start with Maximise Conversions to accumulate conversion data, then test a move to tCPA once the campaign has at least 30–50 conversions per month.[1][10][14] When testing a tCPA or tROAS target change, set the initial target at or near the current observed CPA or ROAS to avoid destabilising the algorithm, then test incremental adjustments of no more than 10–15% at a time.[3][16] Space major Smart Bidding changes at least 60 days apart.[1]
Ad Copy
For Responsive Search Ads, test distinct messaging themes rather than minor wording variations. A theme test might pit benefit-led headlines (“Save $200 on Installation”) against proof-led headlines (“4.9 Stars, 1,200 Reviews”) as a complete angle, rather than swapping a single word. Test value propositions: price, speed, trust signals, differentiation, and offer framing are high-value variables because they directly affect click-through rate and downstream conversion quality.[6][12] Ideal threshold for a reliable ad copy test result is 200 conversions per variation[58], with a minimum run of 14 days and a preferred window of 21–30 days.[58]
Landing Pages
Test message match between the ad headline and the landing-page headline first — this is the highest-leverage single variable in most accounts because misalignment between ad promise and page content directly degrades conversion rate and Quality Score.[12][18] After message match, test conversion friction: form length, CTA prominence, and trust signals. Evaluate landing-page tests with a window that accounts for the full conversion lag, not just initial click-to-lead data.[12][16]
Audiences
Test first-party audiences — customer lists, high-intent site visitors (e.g., visited pricing page), cart abandoners, and CRM-based segments — as bid adjustments or audience targeting layers before testing broader third-party segments. In 2026, first-party data is the primary audience signal of value; third-party cookie deprecation has reduced the reliability of third-party audience targeting.[10][18] For B2B, test intent-based audience layers and value-based segments rather than demographic assumptions, since lead quality matters more than raw volume.[6][12][18]
Worked example
Testing a bidding strategy upgrade from Maximise Conversions to tCPA
- Setup: An Adelaide home security account spending $8,000 per month on Search, running Maximise Conversions for 90 days, now averaging 55 conversions per month at $145 AUD CPA. The team wants to test tCPA to control cost.
- Numbers: 55 conversions per month exceeds the 30–50 per month viability threshold[5]. Minimum duration for a bidding strategy test: 60 days[1]. tCPA target set at $145 (matching observed CPA to avoid destabilisation[3]). 50/50 split: each arm receives ~27–28 conversions per month, or ~55 per arm over 60 days — above the 30 minimum[5]. Start date: 3 August 2026. Ramp-up discarded: 3–10 August 2026[8]. Evaluation window: 11 August – 2 October 2026. Primary metric: CPA. Guardrail: conversion volume must not fall below 20 conversions per month in the treatment arm.
- Decision: Create a campaign experiment draft with tCPA set to $145 AUD, 50/50 split, start 3 August 2026, end 2 October 2026, 95% confidence target[1]. Review results no earlier than 2 October 2026.
- Why: Setting the initial tCPA target at the observed CPA of $145 minimises learning disruption, and the 60-day window provides sufficient conversions per arm for a reliable result.[1][5]
Worked example
Testing a landing-page message-match change for a Search campaign
- Setup: A Canberra recruitment agency spending $6,500 per month on Search, driving traffic to a generic “Find Your Next Role” homepage. The team suspects poor message match is suppressing conversion rate. Current conversion rate from Search: 2.1% (42 conversions per month from ~2,000 clicks). Average CPA: $155 AUD.
- Numbers: Test variable: landing page only — control sends to the existing homepage; treatment sends to a new page with a headline matching the ad (“Find IT Project Manager Roles in Canberra”). 50/50 split. Duration: 28 days[63] minimum for a copy/landing-page test; 42 days preferred to reach 200 conversions per variation[58] (at 50% of 42 monthly conversions = 21 per arm per month; 42 days = ~29 per arm — below the 200 ideal but above the 30 minimum[5], so a 95% confidence result is achievable but may require the full 42 days). Primary metric: conversion rate. Guardrail: CPC must not increase by more than 15% in the treatment arm.
- Decision: Run a 42-day campaign experiment from 5 August 2026 to 16 September 2026, 50/50 split, single variable (landing page URL only), 95% confidence target. Do not edit ad copy, keywords, or bids during the test.
- Why: Message match between ad headline and landing-page headline is the highest-leverage single variable for conversion rate improvement, and isolating it from all other changes is required to attribute any lift correctly.[12][18]
8. Reading Results and Avoiding False Positives
Google’s experiment reporting interface displays performance metrics for the control and treatment arms side by side, with statistical significance indicators. Reading these correctly — and resisting the temptation to act before significance is reached — is the most important skill in experiment management.[11][14]
What the significance indicator means
Google’s system uses a 95% significance bar in experiment reporting, according to practitioner documentation, though Google’s own help pages do not explicitly state a confidence percentage or p-value threshold in the sources available.[14] A result marked as statistically significant means there is a 95% probability that the observed difference is not due to random variation. A result not yet marked as significant means more data is needed — not that the test is failing.
Exclude the ramp-up period
Discard the first 7 days of experiment data before evaluating any metric. Early volatility as the algorithm adjusts to the split is normal and does not represent steady-state performance.[8][11]
Exclude recent days with incomplete attribution
Google’s guidance for value-based bidding experiments explicitly recommends excluding the most recent days where conversion reporting is still incomplete when judging results.[11] For any account with a conversion lag longer than 24 hours, evaluate results at least one full conversion window after the experiment ends, not on the day it closes.
Do not peek and stop early
Evaluating results before the planned end date and stopping the experiment when it looks like a winner — a practice known as “peeking” — inflates the false-positive rate substantially. Commit to the pre-specified end date and confidence threshold before the experiment starts, and hold to both.[1][2][5]
Check the guardrail metric
A treatment that wins on the primary metric but fails the guardrail metric is not a clean win. For example, a treatment that reduces CPA from $150 to $130 but also reduces monthly conversions from 55 to 30 may be optimising for cost at the expense of volume in a way that does not serve the business. Evaluate both metrics before applying the winner.[2][3][5]
Segment results with care
It is valid to segment experiment results by device, audience, or campaign to understand where lift is concentrated, but do not use segmented results as the basis for declaring a winner unless the segmented analysis was pre-specified in the hypothesis. Post-hoc segmentation is a form of multiple testing and inflates false-positive risk.[1][5]
Worked example
Identifying a false positive caused by early stopping
- Setup: A Gold Coast tourism account spending $10,000 per month on Search, running a 28-day ad copy experiment from 1 August 2026 at 50/50. On day 10, the treatment arm shows CPA of $85 AUD versus control CPA of $120 AUD — a 29% improvement. The team considers stopping early.
- Numbers: At day 10, each arm has accumulated approximately 10 days × (50 conversions per month ÷ 30 days) × 50% split = ~8 conversions. 8 conversions per arm is far below the 200-per-variation ideal[58] and below even the 30-minimum[5]. At this volume, a 29% CPA difference has an enormous confidence interval — the result is almost certainly not at 95% significance[1]. The 14-day minimum for ad copy tests has not been met[58].
- Decision: Do not stop the experiment. Continue to the pre-specified end date of 29 August 2026. Re-evaluate results only after 29 August 2026, with at least 14 days of data post-ramp-up and conversion lag cleared.
- Why: Stopping at day 10 with ~8 conversions per arm violates the 14-day minimum[58] and the 30-conversion minimum[5]; the observed 29% CPA improvement cannot be distinguished from random variation at this sample size.[1][2]
9. Applying or Rolling Back a Winner
The decision to apply a winning experiment is not the end of the process — it is the beginning of a controlled rollout. The way a winner is applied determines whether the performance gain is preserved or lost to learning disruption.[15][16]
Applying a winner
Google’s API workflow supports ending or promoting experiments after evaluation. When you promote a winning experiment, Google treats the treatment campaign as a new campaign rather than inheriting the control campaign’s metrics — this is important for reporting continuity and for Smart Bidding, which will enter a new learning phase.[5]
For bidding and target changes, apply winners in small increments rather than making a large single adjustment. If the winning tCPA target is 15% lower than the current target, consider moving in two 7–8% steps spaced 2–3 weeks apart rather than making the full move at once. Give the algorithm 1–2 conversion cycles to stabilise after each incremental change before re-evaluating.[12][16]
Before expanding a winning change to other campaigns, confirm that the campaigns share comparable intent, structure, and conversion type. A winner in a branded Search campaign may not transfer to a non-branded campaign with different user intent and conversion rates.[12][15]
Rolling back a loser
Roll back quickly when the signal is strong — for example, when the treatment arm shows a statistically significant degradation on both the primary metric and the guardrail metric. Roll back by restoring the prior setting, then hold changes steady for at least one full conversion cycle before re-evaluating to get a clean read on recovery.[3][16]
Do not roll back based on early data that has not cleared the ramp-up period and conversion lag. Premature rollback of a genuinely effective change wastes the learning investment and resets the algorithm unnecessarily.[13][16]
When a winner fails the business metric
Roll back if the treatment improves the test metric but degrades the business metric you actually care about. For example, a tROAS target that improves reported ROAS but reduces qualified lead volume or revenue should be rolled back and the hypothesis reconsidered.[10][12][18]
Document the outcome
Whether a test wins, loses, or is inconclusive, document the full record: hypothesis, split, dates, primary metric result, guardrail metric result, statistical significance reached, action taken, and next test planned. This record prevents the same failed hypothesis from being re-tested and allows patterns to emerge across sequential tests.[1][5][9]
Worked example
Applying a winning tCPA experiment incrementally to avoid learning disruption
- Setup: A Hobart accounting software account running a 60-day tCPA experiment concluding 15 October 2026. Control: tCPA $220 AUD, 52 conversions. Treatment: tCPA $220 AUD (test showed the system could hit $195 average CPA at this target), 49 conversions. The experiment reaches 95% significance on CPA improvement from $220 to $195, with conversion volume holding above the 20-per-month guardrail.
- Numbers: Target reduction: from $220 to $195 = $25 reduction = 11.4%. Google recommends incremental changes rather than large jumps[16]; 11.4% is at the upper bound of a conservative 10–15% step[3]. Apply in two steps: Step 1 — lower tCPA to $207 (midpoint, a 6% reduction) on 16 October 2026. Wait 1 conversion cycle (~30 days) = evaluate on 15 November 2026. If stable at or below $207, Step 2 — lower tCPA to $195 on 15 November 2026. Final evaluation: 15 December 2026.
- Decision: Apply winner in two incremental steps: $220 → $207 on 16 October 2026, then $207 → $195 on 15 November 2026 (contingent on Step 1 stability). Do not adjust budget, keywords, or ad copy during this rollout period.
- Why: Large single-step target reductions can destabilise Smart Bidding’s learning cycle; incremental adjustments spaced one conversion cycle apart allow the algorithm to adapt without a full reset.[3][16]
10. Testing with Smart Bidding and Automation
Smart Bidding changes the testing calculus in two important ways. First, any change to a Smart Bidding target, strategy, or campaign structure triggers a new learning phase, during which performance data is noisy and unreliable. Second, Smart Bidding systems optimise dynamically in response to the traffic they observe, meaning that two arms of an experiment may not be truly isolated if they are competing in the same auction pool for the same budget.[3][16]
Experiment, do not just observe
The most defensible way to test a Smart Bidding change is to use a formal campaign experiment — not to simply change the live campaign and observe before and after. Before-and-after analysis without a concurrent control is vulnerable to seasonality, auction shifts, and algorithmic drift that are indistinguishable from the effect of the change.[8][13]
Respect the learning phase
After any Smart Bidding strategy switch, expect a 7–14 day algorithmic learning phase during which performance metrics will fluctuate.[64] Add a further 14-day learning-phase buffer before beginning to evaluate results.[1] Do not judge the experiment during this combined 21–28 day window.
Space major changes 60 days apart
Major Smart Bidding changes made too close together can reset or disrupt the learning cycle cumulatively. Practitioner guidance recommends waiting at least 60 days between major bidding strategy changes and confirming that weekly conversion volume is stable before initiating the next change.[1][3]
AI Max experiments and automation
AI Max experiments are designed specifically for testing AI-powered features within Search campaigns. Their fixed 50/50 split and one-click setup remove the manual draft creation step, but they still require the same evaluation discipline: discard the ramp-up period, wait for the conversion lag to clear, and reach 95% confidence before acting on results.[4][1]
Performance Max asset experiments
PMax asset experiments test creative variables within the Performance Max environment. Given PMax’s fully automated channel allocation, ensure the experiment runs long enough to cover the algorithm’s own optimisation cycle — at least 4–6 weeks per asset group test, with the understanding that testing 10 asset groups sequentially could take 40–60 weeks in total.[4] Prioritise the highest-impact asset groups first.
Worked example
Sequencing a Smart Bidding strategy test to respect the learning phase
- Setup: A Newcastle automotive parts account spending $20,000 per month on Search, currently on Maximise Conversion Value with no tROAS target. The account generates 90 conversions per month with an average order value of $310 AUD and a current ROAS of 1,395% ($279,000 revenue ÷ $20,000 spend). The team wants to test applying a tROAS of 1,400% to control efficiency.
- Numbers: Minimum duration for a bidding strategy test: 60 days[1]. Learning phase after strategy change: 7–14 days[64]. Learning-phase buffer: 14 days[1]. Total do-not-evaluate window: 14 + 14 = 28 days from start. Evaluation window: days 29–60. At 50/50 split, treatment arm receives 45 conversions per month; over the 60-day evaluation window that is ~90 conversions per arm — well above the 30 minimum[5]. tROAS target set at 1,400% (current observed ROAS of 1,395% rounded up by 0.4% to match the test hypothesis of “can tROAS hold ROAS above 1,400% without reducing volume?”). Start date: 3 August 2026. Do-not-evaluate window: 3–31 August 2026. Evaluation window: 1 September – 2 October 2026.
- Decision: Create experiment on 3 August 2026 with tROAS = 1,400%; do not read results until 1 September 2026; require 95% confidence[1] on primary metric (ROAS) and guardrail (conversion volume ≥ 38 per month in treatment arm) before applying winner.
- Why: Evaluating Smart Bidding experiment results during the 7–14 day learning phase[64] produces misleading data; the 28-day do-not-evaluate window ensures the algorithm has stabilised before the formal evaluation begins.[1][64]
11. Common Mistakes to Avoid
The mistakes below are documented consistently across 2026 practitioner sources and Google’s own guidance. Each one has a specific, avoidable cause and a concrete remediation.[1][2][3][5][8][10]
Testing multiple variables simultaneously
Changing bidding strategy, ad headlines, and landing page in the same experiment makes it impossible to attribute any movement in the primary metric to a single cause. The fix is strict: one variable per experiment, documented before launch.[8][1][5]
Running overlapping experiments on the same campaign
Only one experiment can run per campaign at a time.[5] Attempting to run concurrent experiments on the same base campaign is blocked by the platform. Attempting to approximate this by running separate campaigns with different changes simultaneously — without a controlled split — produces unattributable results.
Changing the split mid-experiment
The traffic split percentage cannot be changed after an experiment is set up. If a different split is needed, a new experiment must be created.[2] Attempting to re-weight traffic manually by editing campaign budgets mid-test contaminates the auction split and the results.
Editing campaigns during the experiment
Any change to keywords, bids, ad copy, landing pages, budgets, or audience lists in either the control or the treatment arm during the experiment contaminates the result. Freeze both campaigns for the duration of the test.[8][10][14]
Stopping early on promising early data
Early results are dominated by ramp-up volatility and incomplete attribution. Stopping at day 7 or 10 based on directionally positive data is the leading source of false positives in Google Ads testing. Commit to the pre-specified end date.[1][2][5]
Testing on low-volume campaigns
A campaign generating fewer than 30–50 conversions per month cannot support a reliable experiment conclusion within a practical timeframe.[5] Consolidate campaigns, increase budgets, or wait until volume is sufficient before experimenting.
Failing to account for conversion lag
Reading experiment results on the day the experiment closes — without waiting for the conversion lag to clear — systematically undercounts conversions in the final portion of the test window. Always extend the evaluation window by at least one full conversion lag period after the experiment end date.[11][12]
Testing trivial changes
An experiment that changes a single word in a headline, or adjusts a tCPA target by 2%, is unlikely to generate a signal large enough to be detected above the noise floor within a practical timeframe. Reserve experiments for changes that are hypothesised to move the primary metric by at least 10–15%.[14][7]
Cherry-picking metrics post-hoc
Declaring a winner on a metric that was not the pre-specified primary metric — because the primary metric did not show a significant result — is a form of data dredging. It inflates the false-positive rate in proportion to the number of metrics examined. Specify the primary metric and guardrail metric before the experiment starts and do not change them.[1][3][5]
12. What Changed Recently
The following changes are documented in Google’s own materials and practitioner sources as of late July to August 2026. Senior practitioners should review each item and update their testing workflows accordingly.[2][3][4][5][11]
AI Max experiments now available for Search campaigns
Google has expanded documentation and availability of AI Max experiments — a one-click test that evaluates AI features including search term matching and asset optimisation within existing Search campaigns. The key operational difference from standard campaign experiments is that AI Max experiments split the existing campaign 50/50 automatically, without requiring a campaign copy or draft. This reduces setup time significantly and is the recommended path for testing AI search features in 2026.[4]
Broad Match experiments now provide incremental query insights
Broad Match experiments in Search campaigns have been updated to surface incremental query data — showing which additional search queries the Broad Match variant is capturing versus the control. This reporting improvement gives practitioners a concrete basis for evaluating the incremental reach of Broad Match rather than relying solely on conversion metrics.[3]
Traffic split immutability clarified for all experiment types
Google’s FAQ now explicitly states that for all experiment types — including custom experiments and Demand Gen experiments — the traffic split percentage cannot be changed after setup. Previously, this constraint was documented primarily for specific experiment types. The clarification means practitioners must finalise their split decision at experiment creation, with no ability to rebalance mid-test. If a different split is required, a new experiment must be created from scratch.[2]
Value-based bidding experiment timeline guidance updated
Google’s guidance for value-based bidding experiments now recommends a structured timeline: a ramp-up period, then at least 30 days of uninterrupted testing, followed by evaluation that explicitly excludes the early ramp period and the most recent days where conversion reporting is incomplete.[11] This formalises what was previously implied in general experiment guidance and provides a clearer minimum for bidding strategy tests.
Promoted experiment campaigns are treated as new campaigns
Google’s API documentation confirms that when a treatment campaign is promoted after a successful experiment, it is treated as a new campaign rather than inheriting the control campaign’s historical metrics and Smart Bidding signals.[5] This has implications for campaign history, Quality Score continuity, and the Smart Bidding learning phase that follows promotion. Practitioners should plan for a renewed learning phase of 7–14 days after promoting any winning experiment.[64]
Experiment start date lead time recommended at 3–7 days
Google’s API documentation recommends setting experiment start dates 3–7 days in the future to allow for review and approval processes.[5] This is a platform recommendation rather than a hard constraint, but aligning to it reduces the risk of experiments being delayed by policy review after the intended start date.
Worked example
Updating an existing testing workflow to reflect August 2026 changes
- Setup: A Melbourne digital marketing team managing eight Google Ads accounts, previously running all bidding tests as custom campaign experiments with manually created drafts, using a 50/50 split set at launch and reading results 3 days after experiment close. They need to update their standard operating procedure to reflect the August 2026 platform changes.
- Numbers: Change 1 — AI Max experiments: for any Search campaign test of AI matching or asset features, replace the draft + custom experiment workflow with a one-click AI Max experiment; saves approximately 30–45 minutes of setup time per test across 8 accounts[4]. Change 2 — Split immutability: add a mandatory “split finalisation” checklist step at experiment creation; if a 70/30 split is needed partway through a test, the current experiment must be ended and a new one created — document this in the SOP so it is not done ad hoc[2]. Change 3 — Post-promotion learning phase: after promoting any winning experiment, add a 14-day do-not-evaluate buffer to all reporting dashboards before comparing the promoted campaign’s performance to the pre-experiment baseline[64]. Change 4 — Results evaluation timing: extend the post-close evaluation window from 3 days to a minimum of 14 days (conversion lag) for all accounts with a conversion window longer than 24 hours[11].
- Decision: Update the standard operating procedure by 10 August 2026 to include: AI Max experiment path for Search AI feature tests; split-lock checklist at creation; 14-day post-promotion buffer; and 14-day post-close evaluation delay for all accounts with conversion windows ≥ 48 hours.
- Why: The August 2026 platform changes to split immutability[2], AI Max experiment availability[4], and promoted-campaign treatment[5] each create operational failure points if the existing SOP is not updated before the next experiment cycle.[11]
References
- [1] https://www.growthspreeofficial.com/blogs/best-tricks-and-tips-for-google-ads-experimentat… www.growthspreeofficial.com
- [2] https://creativeadslab.com/google-ads-a-b-testing-2026-growth-hacks/ creativeadslab.com
- [3] https://datadrivengrowthstudio.com/google-ads-2026-experimentation-for-roas-growth/ datadrivengrowthstudio.com
- [4] https://www.digitalapplied.com/blog/performance-max-asset-experiments-2026-test-playbook www.digitalapplied.com
- [5] https://datadrivengrowthstudio.com/google-ads-a-b-testing-success-in-2026/ datadrivengrowthstudio.com
- [6] https://directiveconsulting.com/blog/the-b2b-marketers-guide-to-google-ads-best-practices-… directiveconsulting.com
- [7] https://www.causalfunnel.com/blog/google-ads-a-b-testing-the-2026-step-by-step-playbook-to… www.causalfunnel.com
- [8] https://services.google.com/fh/files/misc/campaign_experiments_two_sheeter.pdf services.google.com
- [9] https://sagum.com/2026/04/02/how-can-i-use-google-ads-experiments-to-test-different-ad-var… sagum.com
- [10] https://almcorp.com/blog/google-ads-campaign-mix-experiments-complete-guide/ almcorp.com
- [11] https://support.google.com/google-ads/answer/10682377?hl=en support.google.com
- [12] https://support.google.com/google-ads/answer/13826584?hl=en support.google.com
- [13] https://business.google.com/us/ad-tools/google-ad-experiments/ business.google.com
- [14] https://support.google.com/google-ads/answer/6154846?hl=en support.google.com
- [15] https://developers.google.com/google-ads/api/docs/experiments/overview developers.google.com
- [16] https://leadsbridge.com/blog/google-ads-best-practices/ leadsbridge.com
- [17] https://ppchero.com/advanced-google-ads-techniques-to-master-in-2026/ ppchero.com
- [18] https://support.google.com/google-ads/thread/413560562/google-ads-best-practices-for-2026?… support.google.com
- [19] https://www.reddit.com/r/googleads/comments/1pdhpd8/what_are_your_google_ad_tips_for_2026/ www.reddit.com
- [20] https://startupscenedaily.com/google-ads-2026-5-steps-to-data-driven-wins/ startupscenedaily.com
- [21] https://aeogrowthstudio.com/google-ads-a-b-testing-2026-experiment-wins/ aeogrowthstudio.com
- [22] https://www.karooya.com/blog/google-ads-experiments-a-b-test-bidding-creatives-campaign-ch… www.karooya.com
- [23] https://dashthis.com/blog/google-ads-best-practices/ dashthis.com
- [24] https://www.searchenginejournal.com/google-ads-bidding-strategies-where-to-spend-your-time… www.searchenginejournal.com
- [25] https://www.youtube.com/watch?v=BmCaZI3JA3s&vl=en www.youtube.com
- [26] https://www.youtube.com/watch?v=NQCLCLR3NQU www.youtube.com
- [27] https://granularmarketing.com/blog/all-available-bid-strategies-in-google-ads-updated-for-… granularmarketing.com
- [28] https://www.youtube.com/watch?v=NRws5-cA_8w www.youtube.com
- [29] https://support.google.com/google-ads/thread/413560562 support.google.com
- [30] https://www.youtube.com/watch?v=LJBwqY4Qr_A www.youtube.com
- [31] https://www.youtube.com/watch?v=hkPl2v2IBXI www.youtube.com
- [32] https://www.reddit.com/r/googleads/comments/1qaoyec/its_2026_whats_one_piece_of_google_ads… www.reddit.com
- [33] https://www.groas.com/post/google-ads-best-practices-2026-rules-that-actually-matter www.groas.com
- [34] https://www.brandingmarketingagency.com/blogs/best-bidding-strategies-for-google-ads-to-dr… www.brandingmarketingagency.com
- [35] https://business.google.com/en-all/accelerate/podcasts/ads-decoded-s1e4/ business.google.com
- [36] https://www.youtube.com/watch?v=Y4KPzjfOjwU www.youtube.com
- [37] https://www.mayple.com/resources/googleadvertising/bidding-strategies-google-ads www.mayple.com
- [38] https://support.google.com/google-ads/answer/17061251?hl=en support.google.com
- [39] https://blog.google/products/ads-commerce/bidding-budgeting-google-marketing-live-2026/ blog.google
- [40] https://twominutereports.com/blog/google-ads-best-practices twominutereports.com
- [41] https://www.wizardcreativelabs.com/post/the-2026-google-ads-guide-campaign-types-bidding-a… www.wizardcreativelabs.com
- [42] https://sandstormdigital.com/2026/05/07/google-ads-bid-strategy-in-2026/ sandstormdigital.com
- [43] https://support.google.com/google-ads/answer/2454137?hl=en support.google.com
- [44] https://support.google.com/google-ads/answer/6318747?hl=en support.google.com
- [45] https://developers.google.com/google-ads/api/docs/experiments/reporting developers.google.com
- [46] https://developers.google.com/google-ads/api/rest/reference/rest/v20/customers.experiments developers.google.com
- [47] https://www.reddit.com/r/PPC/comments/1obeue4/is_7_days_the_minimum_duration_for_google_ad… www.reddit.com
- [48] https://developers.google.com/google-ads/api/docs/change-event?hl=zh-cn developers.google.com
- [49] https://support.google.com/google-ads/answer/6318747?hl=lv-uk support.google.com
- [50] https://support.google.com/google-ads/answer/14147337?hl=en support.google.com
- [51] https://business.google.com/br/ad-tools/google-ad-experiments/ business.google.com
- [52] https://growthmindedmarketing.com/blog/google-ads-experiments/ growthmindedmarketing.com
- [53] https://www.datafeedwatch.com/blog/google-ads-experiments-guide www.datafeedwatch.com
- [54] https://www.linkedin.com/posts/adamjlasky_couple-of-things-i-learned-when-running-google-a… www.linkedin.com
- [55] https://www.youtube.com/watch?v=97F1Voc9tiI www.youtube.com
- [56] https://blog.adnabu.com/google-ads/google-ads-ab-testing/ blog.adnabu.com
- [57] https://www.mbadv.agency/google-ads/optimization-and-testing www.mbadv.agency
- [58] https://paidmediastudio.com/google-ads-a-b-testing-2026s-new-rules-for-marketers/ paidmediastudio.com
- [59] https://almcorp.com/blog/google-ads-experiment-center-guide/ almcorp.com
- [60] https://support.google.com/google-ads/answer/6318731?hl=en support.google.com
- [61] https://almcorp.com/blog/google-shopping-product-data-experiments-2026/ almcorp.com
- [62] https://support.google.com/google-ads/answer/6318742?hl=en support.google.com
- [63] https://www.youtube.com/watch?v=AXkfIvpvszE www.youtube.com
- [64] https://www.searchenginejournal.com/how-to-test-a-new-bid-strategy-in-google-ads/571237/ www.searchenginejournal.com
- [65] https://www.reddit.com/r/adwords/comments/1tpepds/can_google_ads_seriously_not_run_a_4camp… www.reddit.com
- [66] https://www.webtopia.co/blog/google-ads-experiment-center-2026-a-smarter-way-to-run-google… www.webtopia.co
- [67] http://ads-developers.googleblog.com/2026/ ads-developers.googleblog.com
Don’t want to do this yourself?
Get a complete, ready-to-launch Google Ads campaign — keywords validated against real search data, copy written to the best practices on this page, and a proper negative-keyword list — built for your business and delivered within 24 hours.
One-off price · Human-reviewed · No subscription · No access to your Google Ads account
This page is maintained by Sean Cooney at Omologist.com. Content is refreshed every two months using real-time research from authoritative Google Ads sources. Worked examples are illustrative scenarios calculated from published benchmarks, not client results.

