Your Dashboard Said 316%. The Experiment Said 73%.

Every number in your reporting answers the question "which click came before this sale". None of them answers "would the sale have happened anyway". Here is how far apart those two answers have been measured to be, and why more accurate tracking makes the first one better without touching the second.

I have spent nineteen years buying traffic and I have never once had an argument about attribution that was really about attribution. The argument is always about budget, and the budget argument always ends with somebody pointing at a ROAS column as though it settled something.

It does not settle anything, and the reason is structural rather than technical. Your tracker records what happened to people who saw your ad. To know what the ad did, you need to know what would have happened to those same people if they had not seen it. That group does not exist in your data. It cannot, because you showed them the ad.

That missing group is the counterfactual. The distance between what your reporting credits to a campaign and what a randomised experiment measures it caused is what I have started calling the counterfactual gap, and unlike most of the things I write about, this one has been measured properly by people with access to the raw platform data.

Disclosure: ClickerVolt is our product. We aim for fairness in every comparison: we credit competitors where they excel and only highlight genuine gaps. All pricing and features are verified against live sources.

What Attribution Feels Like Against What It Is

Four things a tracker records, and what each one is not FEELS LIKE ACTUALLY IS Click, then purchase The ad made the sale An ordering of events Brand search converts at 40% Our best channel People already sold Retargeting shows 8x ROAS Retargeting is working Warm buyers, re-billed We fixed our tracking Now the numbers are true A cleaner correlation Budget moves toward whatever the model was already going to credit None of the four rows on the left is wrong. Each one is a true statement about sequence being read as a statement about cause.

Every row is accurate as recorded; the error happens in the second column, where an ordering of events becomes a claim about why.

The fourth row is the one I want to sit on, because it is the one this website is otherwise dedicated to. I sell better tracking. Better tracking is worth having. It also does not touch this problem at all, and I would rather say that on my own site than have somebody else say it about me.

The Numbers From the Largest Test I Can Find

In 2019, a team including researchers at Facebook published a comparison of observational advertising measurement against randomised experiments, using 15 US Facebook studies, 500 million user-experiment observations and 1.6 billion impressions. They had something almost nobody has: the randomised answer and the dashboard answer for the same campaigns.

Their headline conclusion is a sentence worth memorising. In half of the studies, the estimated percentage increase in purchase outcomes was off by a factor of three across all methods, and the observational methods generally overestimated.

One study in that set is the clearest illustration of the counterfactual gap I have ever seen published.

316% purchase lift from the naive comparison of exposed against unexposed users
222% purchase lift after exact matching on age and gender
73% purchase lift measured by the randomised experiment

Look at the middle number. Somebody did the responsible thing. They did not just compare people who saw the ad against people who did not, they matched on demographics first, which is exactly the kind of rigour that gets you nodded at in a meeting. It cut the error roughly in half and still overstated the truth by a factor of three.

Same campaign, same data, three ways of asking Marketing Science, 2019, study 4 of 15. Purchase lift. NAIVE MATCHED RANDOMISED +316% +222% +73% Saw the ad vs did not Exact match on age and gender Assignment decided by coin flip Overstated 4.3x Overstated 3.0x The measurement CLEANER DATA MOVED IT. ONLY RANDOMISATION LANDED IT. The counterfactual gap is the distance between the middle box and the right box. It is not tracking error. Both estimates used the same, correct, complete data. It is selection: the people who saw the ad were already more likely to buy. Across all 15 studies, half were off by a factor of three.

Nothing between the left box and the middle box was a data quality improvement you could buy; the remaining error was in the question, not the inputs.

The Cleanest Version of the Same Story

The eBay experiment is older and blunter, and it is the one I would use if I had thirty seconds to make this point to a sceptic.

In 2015, researchers published a field experiment in Econometrica where eBay switched off brand-keyword advertising on Yahoo and MSN while holding Google as a control. When the paid brand ads stopped, 99.5% of the forgone paid click traffic came back through natural search instead. The authors describe that as a lower bound on retention.

The same paper reported eBay's non-brand paid search at an observational return of 1,632% once time and geography were controlled for, against an experimental return of negative 63%.

99.5% of eBay's forgone paid brand clicks recaptured by natural search when brand ads stopped

That is not an argument that paid search does not work. The same paper found new and infrequent users did respond to ads. The finding is subtler and more expensive: the responsive people were a minority, the unresponsive frequent buyers absorbed most of the spend, and the reporting could not tell those two groups apart because both of them clicked.

What This Does to an Affiliate Specifically

Affiliates run exactly the campaign shapes with the widest counterfactual gap and the most flattering dashboards.

Retargeting is the obvious one. It shows the best ROAS in every account I have ever looked at, and the honest measured numbers are far smaller. A ghost-ad controlled experiment on display retargeting found 17.2% more site visits and 10.5% more purchases. A separate randomised study found retargeting caused 14.6% more users to return within four weeks. Those are real effects, worth having, and they are nowhere near an 8x return.

Brand and brand-adjacent search is the other one, and eBay covered it.

Then there is a second error stacked on top, and this one is arithmetic rather than causal. If your refunds, rebills and chargebacks never make it back to the ad platform, the attributed number is inflated before anyone even asks whether the ad caused anything. Two independent overstatements, pointing the same way, multiplying.

Which campaign has the widest gap? Test that one first. Pick a campaign Does it target people who already know you? YES Widest gap. Test this first. Brand search, retargeting, email-matched audiences NO Do refunds and rebills reach the platform? NO Fix the arithmetic before the causal test A holdout on inflated inputs measures the wrong baseline YES Can you go dark in a matched set of geos? NO Use the platform's own A/B holdout instead YES Run the geo holdout. This is the measurement.

The order matters: an incrementality test run on top of an inflated conversion count measures a gap you introduced yourself.

Why This Is Surfacing Now

Because the platforms have started conceding it in public, which they did not use to.

Meta launched Incremental Attribution around April 2025 and now offers it in Ads Manager as both an optimisation setting and a reporting column. Meta's own newsroom said in January 2026 that its Q4 model rollout drove a 24% increase in incremental conversions against the standard attribution model, and that the product reached a multi-billion-dollar annual run rate seven months after launch. Read that carefully. Meta is now selling a distinct product on the basis that its standard attribution and its incremental attribution are different quantities.

In June 2026, a TikTok research team published a paper whose title is the whole argument: Attributed, But Not Incremental. It introduced a name I have adopted, the cannibalization rate, defined as the fraction of nominal paid-attributed conversions that is not truly incremental. Their stated concern is that paid-attributed conversions systematically overstate true growth when the paid channel overlaps with organic demand, brand-driven traffic or another acquisition channel. After reallocating budget using a corrected signal across multiple global markets, the measured overall cannibalization rate fell by roughly 15 percentage points. The paper does not publish the absolute level, only that change, and I am not going to turn one into the other.

Google, meanwhile, dropped the price. Its documentation says an experiment that once cost upwards of $100,000 can now be run for $5,000, which it attributes to moving Conversion Lift from Frequentist to Bayesian methodology and using historical campaign data as priors.

The catch is access. Google Ads Help still states plainly that Conversion Lift is not available for all accounts and that you should contact your account representative. Meta's Conversion Lift documentation routes you to a rep for minimums and does not publish a figure. Cheaper, and still not open.

What Good Looks Like

If you cannot get a rep, and as an affiliate you probably cannot, the do-it-yourself version is a geo holdout.

Pick a set of geographies whose historical revenue moves together. Keep the campaign running in one set and switch it off entirely in the other. Leave it off long enough that the effect and the noise separate, which is weeks rather than days. Then model what the dark set would have done from its own history and the control's movement, and take the difference.

The tooling is free and open source. Google's CausalImpact is actively maintained, with version 1.4.1 published in September 2025. Google's Meridian GeoX, announced in May 2026, is publisher-agnostic and built for exactly this, with go-dark and holdback designs and results that feed back into a media mix model. Meta's GeoLift is the package everyone cites and its most recent release is dated 19 May 2023, which you should know before you build a practice on it.

Two things will decide whether the result is worth anything. Contamination, meaning people exposed in a live market who convert in a dark one, which is why matched markets are chosen algorithmically rather than by intuition. And the completeness of your conversion data, because if refunds and rebills are missing from both sides of the comparison, you have measured the lift in gross bookings rather than in money.

So What Do You Do About It

The counterfactual gap is not a bug in your tracker and it is not something a better tracker closes. It is the distance between the question your reporting can answer and the question your budget decision requires, and the only instrument that crosses it is an experiment you have to pay for in withheld revenue.

I will be direct about my own position. ClickerVolt does not measure incrementality. No tracker does. What it does is make the inputs to a lift test trustworthy: fifteen Meta customer-information parameters so the platform-side numbers are not degraded by weak matching, Refund Sync so a reversal reaches Google, Meta and TikTok on the day it posts rather than never, and unlimited retention so the pre-period you compare against is still there when you need it. You can see how that part works here.

The free version costs you nothing this week. Take the campaign with your highest reported ROAS. Write down, on paper, what you think total revenue would do if it went dark for fourteen days. Then write down what evidence you have for that number. If the only evidence is the campaign's own reporting, you have just identified the campaign most worth testing, and you have done it without buying anything.

Ready to Track Smarter?
Start for free — every feature, no credit card, no trial countdown.
Try ClickerVolt Free →
500 events/month free · All features included · No credit card