What Attribution Feels Like Against What It Is
Every row is accurate as recorded; the error happens in the second column, where an ordering of events becomes a claim about why.
The fourth row is the one I want to sit on, because it is the one this website is otherwise dedicated to. I sell better tracking. Better tracking is worth having. It also does not touch this problem at all, and I would rather say that on my own site than have somebody else say it about me.
The Numbers From the Largest Test I Can Find
In 2019, a team including researchers at Facebook published a comparison of observational advertising measurement against randomised experiments, using 15 US Facebook studies, 500 million user-experiment observations and 1.6 billion impressions. They had something almost nobody has: the randomised answer and the dashboard answer for the same campaigns.
Their headline conclusion is a sentence worth memorising. In half of the studies, the estimated percentage increase in purchase outcomes was off by a factor of three across all methods, and the observational methods generally overestimated.
One study in that set is the clearest illustration of the counterfactual gap I have ever seen published.
Look at the middle number. Somebody did the responsible thing. They did not just compare people who saw the ad against people who did not, they matched on demographics first, which is exactly the kind of rigour that gets you nodded at in a meeting. It cut the error roughly in half and still overstated the truth by a factor of three.
Nothing between the left box and the middle box was a data quality improvement you could buy; the remaining error was in the question, not the inputs.
The Cleanest Version of the Same Story
The eBay experiment is older and blunter, and it is the one I would use if I had thirty seconds to make this point to a sceptic.
In 2015, researchers published a field experiment in Econometrica where eBay switched off brand-keyword advertising on Yahoo and MSN while holding Google as a control. When the paid brand ads stopped, 99.5% of the forgone paid click traffic came back through natural search instead. The authors describe that as a lower bound on retention.
The same paper reported eBay's non-brand paid search at an observational return of 1,632% once time and geography were controlled for, against an experimental return of negative 63%.
That is not an argument that paid search does not work. The same paper found new and infrequent users did respond to ads. The finding is subtler and more expensive: the responsive people were a minority, the unresponsive frequent buyers absorbed most of the spend, and the reporting could not tell those two groups apart because both of them clicked.
What This Does to an Affiliate Specifically
Affiliates run exactly the campaign shapes with the widest counterfactual gap and the most flattering dashboards.
Retargeting is the obvious one. It shows the best ROAS in every account I have ever looked at, and the honest measured numbers are far smaller. A ghost-ad controlled experiment on display retargeting found 17.2% more site visits and 10.5% more purchases. A separate randomised study found retargeting caused 14.6% more users to return within four weeks. Those are real effects, worth having, and they are nowhere near an 8x return.
Brand and brand-adjacent search is the other one, and eBay covered it.
Then there is a second error stacked on top, and this one is arithmetic rather than causal. If your refunds, rebills and chargebacks never make it back to the ad platform, the attributed number is inflated before anyone even asks whether the ad caused anything. Two independent overstatements, pointing the same way, multiplying.
The order matters: an incrementality test run on top of an inflated conversion count measures a gap you introduced yourself.
Why This Is Surfacing Now
Because the platforms have started conceding it in public, which they did not use to.
Meta launched Incremental Attribution around April 2025 and now offers it in Ads Manager as both an optimisation setting and a reporting column. Meta's own newsroom said in January 2026 that its Q4 model rollout drove a 24% increase in incremental conversions against the standard attribution model, and that the product reached a multi-billion-dollar annual run rate seven months after launch. Read that carefully. Meta is now selling a distinct product on the basis that its standard attribution and its incremental attribution are different quantities.
In June 2026, a TikTok research team published a paper whose title is the whole argument: Attributed, But Not Incremental. It introduced a name I have adopted, the cannibalization rate, defined as the fraction of nominal paid-attributed conversions that is not truly incremental. Their stated concern is that paid-attributed conversions systematically overstate true growth when the paid channel overlaps with organic demand, brand-driven traffic or another acquisition channel. After reallocating budget using a corrected signal across multiple global markets, the measured overall cannibalization rate fell by roughly 15 percentage points. The paper does not publish the absolute level, only that change, and I am not going to turn one into the other.
Google, meanwhile, dropped the price. Its documentation says an experiment that once cost upwards of $100,000 can now be run for $5,000, which it attributes to moving Conversion Lift from Frequentist to Bayesian methodology and using historical campaign data as priors.
The catch is access. Google Ads Help still states plainly that Conversion Lift is not available for all accounts and that you should contact your account representative. Meta's Conversion Lift documentation routes you to a rep for minimums and does not publish a figure. Cheaper, and still not open.
What Good Looks Like
If you cannot get a rep, and as an affiliate you probably cannot, the do-it-yourself version is a geo holdout.
Pick a set of geographies whose historical revenue moves together. Keep the campaign running in one set and switch it off entirely in the other. Leave it off long enough that the effect and the noise separate, which is weeks rather than days. Then model what the dark set would have done from its own history and the control's movement, and take the difference.
The tooling is free and open source. Google's CausalImpact is actively maintained, with version 1.4.1 published in September 2025. Google's Meridian GeoX, announced in May 2026, is publisher-agnostic and built for exactly this, with go-dark and holdback designs and results that feed back into a media mix model. Meta's GeoLift is the package everyone cites and its most recent release is dated 19 May 2023, which you should know before you build a practice on it.
Two things will decide whether the result is worth anything. Contamination, meaning people exposed in a live market who convert in a dark one, which is why matched markets are chosen algorithmically rather than by intuition. And the completeness of your conversion data, because if refunds and rebills are missing from both sides of the comparison, you have measured the lift in gross bookings rather than in money.
So What Do You Do About It
The counterfactual gap is not a bug in your tracker and it is not something a better tracker closes. It is the distance between the question your reporting can answer and the question your budget decision requires, and the only instrument that crosses it is an experiment you have to pay for in withheld revenue.
I will be direct about my own position. ClickerVolt does not measure incrementality. No tracker does. What it does is make the inputs to a lift test trustworthy: fifteen Meta customer-information parameters so the platform-side numbers are not degraded by weak matching, Refund Sync so a reversal reaches Google, Meta and TikTok on the day it posts rather than never, and unlimited retention so the pre-period you compare against is still there when you need it. You can see how that part works here.
The free version costs you nothing this week. Take the campaign with your highest reported ROAS. Write down, on paper, what you think total revenue would do if it went dark for fourteen days. Then write down what evidence you have for that number. If the only evidence is the campaign's own reporting, you have just identified the campaign most worth testing, and you have done it without buying anything.
