How to Run a Holdout Test on Affiliate Campaigns

A holdout test is the only measurement in your entire stack that tells you what an ad caused rather than what it preceded. It is also the only one that costs you revenue to perform, which is why almost nobody runs one.

This is the version an affiliate can actually execute. No account representative, no enterprise measurement vendor, no six-figure budget. A geo split, a fixed number of weeks of discipline, and an open-source read-out.

Before the steps, one number to justify the effort. In a comparison of observational advertising measurement against randomised experiments across 15 Facebook studies, half the studies were off by a factor of three, with the observational estimates generally too high. In one, the randomised answer was 73% purchase lift while the naive read said 316%, and matching on age and gender still said 222%. You cannot shortcut your way to the right answer with cleaner data. You have to withhold the ad from somebody.

Disclosure: ClickerVolt is our product. We aim for fairness in every comparison: we credit competitors where they excel and only highlight genuine gaps. All pricing and features are verified against live sources.

Step 1: Pick the Campaign With the Most to Hide

Do not start with the campaign you are least sure about. Start with the one whose reported numbers are so good that nobody has questioned them.

In practice that means whichever of these you run:

  • Branded or brand-adjacent search
  • Retargeting and any warm-audience remarketing
  • Email-matched or customer-list audiences
  • Any campaign whose audience definition includes people who have already visited your page

These are the campaigns where the ad most plausibly reached somebody who was going to buy anyway. When eBay switched off brand-keyword advertising in a controlled experiment, 99.5% of the forgone paid click traffic came back through natural search, and the authors called that a lower bound. Cold prospecting has its own problems, but credit-without-cause is not usually the main one.

Pick exactly one campaign. Testing two things at once produces a result you cannot attribute to either.

Step 2: Fix the Arithmetic Before You Test the Causation

This step is boring and skipping it will waste the whole exercise.

A holdout compares revenue in live markets against revenue in dark markets. If your refunds, chargebacks and rebills are not reaching your conversion data, then both sides of that comparison are measuring gross bookings rather than money, and your measured lift will be a lift in a number that nobody banks.

Before you start, confirm three things:

  1. Refunds and chargebacks reduce the revenue figure you will use as the outcome metric.
  2. Rebills and upsells are included in it, attributed back to the original click.
  3. The outcome metric you will measure is available at the geography level, daily.

That last one catches people. If your revenue only exists as a network-level monthly statement with no geography attached, you cannot run a geo test, and you need to fix that first.

Step 3: Choose the Design

Three shapes, all of them named in Google's own open-source geo framework.

Go dark. Switch the campaign off entirely in the test markets. Strongest signal, largest cost, clearest result. This is what you want for a warm-audience campaign you suspect of being non-incremental, because the whole question is what happens when it stops.

Holdback. Withhold from a subset while the rest continues. Same logic as go dark, applied to a smaller share of your footprint.

Heavy up. Increase spend in test markets rather than removing it. Useful when you want to know the return on the next dollar rather than on the whole campaign, and it does not cost you revenue while it runs. Weaker for answering "should this campaign exist".

For a first test, go dark. The answer is unambiguous and you will learn more from it than from a subtle one.

Step 4: Choose Matched Markets, Not Markets That Feel Similar

This is where self-run tests fail most often, so I am going to be specific.

You need two groups of geographies whose historical revenue moves together, so that the control group can stand in for what the test group would have done. Choosing them by intuition, or by picking states that seem comparable, is not good enough.

Google's own geo lift product does not use states or metros as they exist administratively. It builds its experimental units, Google Marketing Areas, by spectral clustering, with an explicit model for people travelling between areas. That is not corporate excess. It is a direct response to the single biggest threat to a geo test.

The shape of one go-dark geo test PRE-PERIOD TEST WINDOW READ-OUT 8 to 12 weeks minimum 21 days dark, nothing else changes Plus 7 days of carryover Correlate daily revenue by geo Campaign off in test geos only Model the counterfactual Split into matched pairs No creative or budget changes anywhere Report an interval, not a point Write the end date down first Do not peek and do not stop early Convert to iROAS and decide THE THREAT IS CONTAMINATION Someone exposed in a live market who converts in a dark market drags measured lift down. Google builds its experimental units by spectral clustering with a travel model for exactly this. Picking states that feel similar is the most common way a self-run geo test dies. Pair markets on historical daily revenue correlation, and keep commuter-adjacent areas on the same side.

The pre-period is not preparation, it is data: without a clean history to correlate on, there is no way to construct the counterfactual.

The practical version, if you are doing this by hand:

  • Pull daily revenue by geography for the last 8 to 12 weeks at minimum. Longer is better.
  • Correlate each geography's daily series against every other one.
  • Form pairs with high correlation and similar volume, then assign one of each pair to test and one to control.
  • Keep geographies that share a commuting or media footprint on the same side of the split.
  • Exclude anything with a volume too small to produce a stable daily series. A geography with three conversions a week adds noise, not signal.

Step 5: Size It Honestly Before You Start

Google's geo lift setup screen does something worth copying even if you never get access to it. Before you commit, it shows a feasibility rating of High, Medium or Low, and a minimum detectable iROAS. Google's own advice is not to proceed when feasibility comes back Low.

That is the discipline to import: decide, in advance, what size of effect your test could actually detect. If the honest answer is that you could only detect a 60% swing and you are trying to distinguish a 1.2x return from a 2.0x return, the test will not answer your question and running it is worse than not running it, because you will get a number and treat it as evidence.

Scale matters here too. Google's documentation now says an experiment that once cost upwards of $100,000 can be run for $5,000, which it attributes to moving Conversion Lift from Frequentist to Bayesian methodology, using historical campaign data as priors. That is a real reduction, and it is the reason a smaller advertiser can now think about this at all. It is also worth saying in the same breath that Google Ads Help still states plainly that Conversion Lift is not available for all accounts and that you should contact your account representative. Cheaper, still gated.

Step 6: Freeze Everything Else

For the duration of the test, in both test and control markets:

  • No creative changes
  • No budget changes to the campaign under test or to any campaign that touches the same audience
  • No landing page tests
  • No new offers introduced or retired
  • No price changes or promotions

And check the calendar before you set the dates. A test window that straddles a major shopping event, a holiday period or a known seasonal spike in one region and not another produces a number you will not be able to defend.

Step 7: Run It Long Enough, and Do Not Stop Early

Three weeks dark is a reasonable first target for most affiliate campaigns, followed by about a week of carryover observation, because demand does not stop the instant an ad does.

Write the end date down before you begin and treat it as fixed. The single most common way a well-designed test produces a meaningless result is that somebody watched revenue dip in week one, panicked, and turned the campaign back on. Peeking at an experiment and stopping when you like the answer is not measurement, it is a slower way of confirming what you already believed.

Step 8: Read It Out

You are trying to estimate what the dark markets would have earned if the campaign had kept running, then subtract that from what they actually earned.

Three tools, in the order I would reach for them:

CausalImpact. Google's Bayesian structural time-series package for R, Apache licensed and actively maintained, with version 1.4.1 published in September 2025. You give it the outcome series for the test markets, the control markets as covariates, and the intervention date. It returns the estimated counterfactual and a credible interval. This is the most accessible option and it is the one I would start with.

Meridian GeoX. Google's open-source, publisher-agnostic geo incrementality framework, announced in May 2026. It supports holdback, go-dark and heavy-up designs natively, handles multiple test cells against a common control, and produces results that feed back into a media mix model as priors. More setup than CausalImpact, more capability.

GeoLift. Meta's synthetic control package, MIT licensed, and the one you will see recommended most often. Its repository is public and marked active, and its most recent release, v2.6.06, is dated 19 May 2023. Installation needs remotes and augsynth, and its own release notes record working around a dependency being archived from CRAN. It still works and plenty of people use it. Know the release date before you build a practice on it.

Whichever you use, expect an interval rather than a single number. Google's own move to Bayesian methodology reports 80% credible intervals for the same reason: a point estimate from a geo experiment implies a precision the design does not have.

Step 9: Convert to iROAS and Actually Decide

Incremental ROAS is incremental revenue divided by media spend. That is Google's own definition and it is the number you carry out of the test.

The decision rule is simpler than people make it:

  • If iROAS comfortably clears your required return, the campaign is earning its budget and you can stop arguing about it.
  • If iROAS is near or below 1, you are buying conversions you would have received anyway. Cut the budget rather than the campaign, and retest at the lower level.
  • If the interval is so wide that it spans both answers, the test was underpowered. That is a design result, not a finding, and you should not act on it in either direction.

Write the result down with the dates, the markets and the method. In six months somebody will ask why the retargeting budget is what it is, and the answer being "we tested it in March and here is the interval" is worth more than any dashboard screenshot.

If You Cannot Split by Geography

Some affiliates cannot. Your traffic may be too concentrated, your offer may be single-country and small, or your platform may not let you segment cleanly.

The fallback is the platform's own self-serve holdout. Meta's A/B testing inside Experiments supports up to five variants and reports once there are at least 100 conversion events per strategy tested, which is a volume gate rather than a spend gate. That is a genuine randomised comparison and it is available without a representative.

The rep-gated products are the deeper ones. Meta's Conversion Lift documentation directs you to your account representative to find out which tests you qualify for and what minimums apply, and publishes no dollar figure. You will find confident spend minimums quoted on agency blogs. I checked several and they contradict each other, none of them traces to a Meta-published page, so I am not repeating any of them. Ask your rep, if you have one.

One thing worth knowing if you do get access: Meta requires the Conversions API for Conversion Lift tests measuring standard web events started on or after 1 October 2021. Server-side is a prerequisite, not an optimisation.

What This Costs and Why It Is Still Worth It

A go-dark test costs you the incremental revenue of the campaign in the test markets for three weeks. If the campaign turns out to be highly incremental, that is a real, measurable loss, and you should budget for it as the price of the information.

The alternative is continuing to allocate budget on a number that, in the largest published comparison I can find, was wrong by a factor of three in half of the cases.

Where a Tracker Fits, and Where It Does Not

I should be clear about this since I sell one. ClickerVolt does not run holdout tests. No tracker does, and any tool marketing incrementality that is not withholding ads from somebody is doing modelling and calling it measurement.

What a tracker decides is whether the inputs to your test are worth analysing. Refunds and chargebacks reaching your revenue figure rather than sitting in a report, so you measure lift in money rather than in bookings. History still being there when you need a twelve-week pre-period. Conversions attached to a geography at all. ClickerVolt fires the Google RETRACT, the Meta reversal and the TikTok CancelOrder the day a refund posts, and keeps unlimited history on every plan including the free one, which is the part of this exercise it can genuinely help with. You can see how that works here.

Then go and withhold an ad from somebody. That is the part no product does for you.

Ready to Track Smarter?
Start for free — every feature, no credit card, no trial countdown.
Try ClickerVolt Free →
500 events/month free · All features included · No credit card