Is That Channel Incremental, or Would It Have Converted Anyway?
Your best-performing channel usually looks best because of where it sits in the journey, not what it does. Retargeting, branded search, and email to logged-in users sit closest to the purchase, so attribution credits them with conversions that would have happened anyway. That's why pouring budget in makes ROAS fall instead of scale: you're paying a higher toll on a finite pool of already-decided buyers. Attribution measures involvement, not causation. The only tool that separates the two is a holdout test. Attribution shortlists what to test; holdouts tell you what's true.
The channel that wins the report is often the one that did the least
Ask most marketers to name their best channel and they'll point at whatever has the highest ROAS in the dashboard. Usually that's retargeting, branded search, or email to logged-in users. Then they try to scale it, and the number falls apart. The ROAS that was 8x at a small budget slides to 4x, then to 3x, as spend goes up.
The usual explanation is that the channel "hit diminishing returns." That's true but it hides the mechanism. The channel looked good in the first place because of where it sits in the journey, not because of what it does. The closer a touchpoint sits to the moment of purchase, the more conversions get credited to it, and the less of any purchase it actually caused. Retargeting, branded search, and lifecycle email don't create buyers. They intercept people who already decided, stand in the doorway, and collect the receipt.
This is proximity bias, and it's the single most common reason a "winning" channel refuses to scale.
Why proximity inflates credit
Attribution assigns credit by presence. A channel earns a slice of a conversion by being on the path, inside the lookback window, close to the moment of sale. Nothing in that logic asks whether the channel changed the outcome. It only asks whether the channel was there.
So the channels that sit late in the journey collect the most credit almost automatically. By the time someone searches your brand name or clicks the retargeting ad chasing the cart they already loaded, they've decided. The late touch is a witness to the decision, not the cause of it. Attribution can't tell the difference, so it hands the witness the credit.
That's the core claim, and it's worth stating plainly: attribution measures involvement, not causation. Every model in the standard toolkit, from last-click to data-driven multi-touch, reads correlation on the observed path and reports it as contribution. None of them run the counterfactual that would tell you what happened without the touch.
THE MECHANISM IN ONE LINE
A channel that sits close to the purchase gets credited for conversions it merely witnessed. The report calls this "performance." It's really proximity.
The test for whether a channel is a witness or a cause is simple to state and hard to fake: turn it off for a random slice of users and see whether conversions actually fall. If they don't, the credit was never earned.
Why the toll can't scale
Here's why a proximity-biased channel breaks the moment you fund it like a growth engine. It harvests a finite pool of already-decided people. When you double the retargeting budget, you don't buy more buyers. You re-reach the same finite pool at a higher clearing price, and you reach a few more people who would have converted anyway. Incremental conversions per dollar drop while the platform-reported total barely moves.
So the ROAS falls, and the fall gets blamed on the channel "saturating." What actually happened is that the measurement caught up with reality. The 8x had always been a reading of how close the channel sat to the money, and that stops flattering you once you push real budget through it.
What the experiments actually found
This isn't a thought experiment. Some of the largest field studies in advertising exist precisely to measure the gap between what a channel gets credited with and what it causes.
eBay turned off branded search and almost nothing happened. Blake, Nosko and Tadelis ran large-scale experiments at eBay that shut off paid search. For branded keywords, roughly 99.5 percent of the traffic the ads would have delivered arrived at the site anyway through organic search. Brand-keyword ads showed no measurable short-term causal benefit. For non-brand keywords, the effect on frequent shoppers was a statistically insignificant 0.66 percent, and returns to non-brand paid search were negative once you priced in the clicks from users who would have come regardless (Econometrica, 2015).
Uber turned off two-thirds of its ad budget and installs didn't move. Kevin Frisch, Uber's former head of performance marketing, described switching off performance spend: the team turned off 100 million dollars of a 150 million dollar annual budget and saw basically no change in the number of rider app installs. The installs they thought had come through paid channels suddenly showed up as organic. The paid channels had been attributing installs to themselves that were happening anyway. That's proximity bias in production, at nine figures.
Observational models routinely overstate ad effects against the experimental ground truth. Across 15 large advertising experiments at Facebook (more than 500 million user-experiment observations, 1.6 billion impressions), Gordon and colleagues compared randomized lift against the observational attribution methods most marketers actually use. The observational methods frequently failed to recover the true effect, sometimes overstating it several-fold and sometimes getting the sign wrong (Marketing Science, 2019). The gap isn't a rounding error. It's the difference between "scale this" and "cut this."
EBAY BRANDED SEARCH: ATTRIBUTED VS CAUSAL
Question: how much of the traffic that branded-keyword ads take credit for would arrive without them?
- Credit for the branded clicks
- A clean, high ROAS line in the report
- A channel that looked untouchable
The reason platform numbers get believed at all is that individual sales are extremely noisy. Lewis and Rao looked at 25 large field experiments (about 2.8 million dollars in ad spend, most reaching millions of users) and found the median confidence interval on advertising ROI was over 100 percentage points wide (Quarterly Journal of Economics, 2015). Per-person sales have so much variance that a real, profitable ad effect is tiny next to the noise. Observational attribution feels precise because it manufactures a clean number by ignoring the noise a true experiment has to fight through. The clean number is the tell, not the reassurance.
The formal name for the underlying problem is selection. Johnson, Lewis and Nubbemeyer put it exactly: the people who see a retargeting ad are "a highly selected group that may have purchased anyway." Measuring what the ad caused requires knowing how those exposed users would have behaved unexposed, which means a randomized control counterpart, not a lookback window (Journal of Marketing Research, 2017). Their "ghost ads" method exists to separate the two things a lookback window fuses together: the selection effect (who got targeted) and the ad effect (what the ad changed). Attribution reports the sum of both and labels it contribution.
The division of labor: attribution proposes, holdouts dispose
None of this means attribution is useless or that you should switch it off. It means attribution has a narrower job than the dashboard implies. Attribution is a cheap, always-on radar. It's very good at surfacing "this channel touches a lot of conversions," which is exactly the input you need to decide what's worth the expense of a real test. What it cannot do is tell you whether the channel made those conversions happen.
So the operating model isn't attribution versus incrementality. It's a division of labor. Attribution shortlists the hypotheses. A holdout is the referee that promotes a hypothesis to fact. You turn the channel off for a random slice of users, watch whether conversions actually fall, and only then treat the attributed number as real.
If you want the reconciliation math for how attribution, MMM and incrementality fit together into one defensible number, that's covered in MTA, MMM and incrementality: triangulating the truth. This article is about one thing: knowing which channel to point the referee at first. Point it at your highest-ROAS proximity channel, because that's where the gap between credited and caused is usually widest.
Three ways this shows up
A WORKED EXAMPLE
Retargeting that "scaled badly." A DTC brand runs retargeting at $10k/mo and the platform reports 900 conversions, a 9× ROAS. They scale to $30k/mo and the platform now reports 1,350 conversions, a 4.5× ROAS. By $50k the number is 3×. The tell is that incremental conversions per dollar cratered while total conversions barely grew.
A 50/50 user holdout at the $10k level shows the truth. Only about 120 of the 900 conversions disappear when retargeting is off. The causal number is ~120, the real ROAS is closer to 1.2×, and it was never 9×. The channel was harvesting a fixed pool of already-decided carts the whole time, and scaling only exposed it.
Email to logged-in users. An e-commerce team sees abandon-cart and post-purchase emails credited with 25% of attributed revenue, and leadership wants to hire two more lifecycle marketers. A holdout suppresses the abandon-cart email for a random 20% of triggered users. Revenue in the holdout is 3% lower, not 25%. Most of those buyers came back regardless. The channel is worth keeping (3% real lift on near-zero marginal cost is excellent ROI) but it cannot grow the business, because it only ever touches demand that already exists.
The branded-search version is the eBay result in miniature. A B2B SaaS spends $8k/mo bidding on its own brand name, and last-click credits it with 40 percent of demos booked. It looks untouchable. Run a two-week paused-brand holdout in half the geos, matched to the other half, and demos in the paused geos fall by about 2 percent, not 40. When the ad is gone, the top organic result is also them, and people click it for free. The $8k was buying a click they already owned. This is the safest test to run first, because it's the most repeatable finding in the whole literature.
Here's what a matched geo-holdout comparison looks like as a query. The point is to compare paused geos against a matched control window, not to compare a channel against itself over time:
-- Paused-brand-search geo holdout: did demos actually fall?
SELECT
g.cohort, -- 'paused' or 'control'
COUNT(DISTINCT d.demo_id) AS demos,
COUNT(DISTINCT g.geo_id) AS geos,
ROUND(
COUNT(DISTINCT d.demo_id) * 1.0
/ NULLIF(COUNT(DISTINCT g.geo_id), 0), 2) AS demos_per_geo
FROM geo_cohorts g
LEFT JOIN demos d
ON d.geo_id = g.geo_id
AND d.created_at BETWEEN g.test_start AND g.test_end
GROUP BY g.cohort
ORDER BY g.cohort;
-- Read the gap in demos_per_geo between paused and control.
-- A 40% attributed channel that only moves this 2% was never causal.
Where this argument stops
This audience is senior, so the overclaims matter as much as the claim. Four caveats keep the point honest.
Proximity channels are not worthless. Branded search, retargeting and lifecycle email frequently have genuinely positive, if small, incremental ROI and near-zero marginal cost. The claim is that they can't scale acquisition, not that you should turn them off. A channel that returns 3 percent real lift on almost no marginal spend is a good channel to run. It's a bad channel to fund like a growth engine.
MTA does not measure incrementality, and this audience knows it. Multi-touch attribution, including the "data-driven" and algorithmic flavors, redistributes credit across observed touchpoints using correlational rules. It has no counterfactual. It cannot tell you what would have happened without the touch. MTA's honest job is shortlisting hypotheses for testing. Treating it as a causal tool is the fastest way to lose a measurement-literate reader.
Holdouts have real costs and failure modes. They forgo revenue on the held-out slice. They need adequate statistical power, and Lewis and Rao show that informative ad experiments can require millions of person-weeks of exposure, which small brands often can't muster. In retargeting the control counterpart is genuinely hard to build, which is the whole reason ghost ads were invented. And a holdout goes stale as seasonality and creative shift. A holdout is a referee you bring in for a specific match, not a permanent installation.
Proximity bias is a strong tendency, not a law. A late touch can be causal. A retargeting ad can surface a genuinely forgotten cart. A branded-search ad can defend against a competitor conquesting your name. Late doesn't equal fake. The working rule is that a late touch is presumed non-causal until a holdout proves otherwise, which sets the direction of the burden of proof and puts that burden on the channel with the suspiciously clean ROAS.
The through-line
The channel with the best ROAS is usually the one you understand least, because its number only ever measured where it sat in the journey, not what it caused. The dashboard rewards proximity and calls it performance, which is exactly why the winner refuses to scale: extra budget buys a higher toll on buyers you already had.
Attribution still earns its keep as the radar that tells you which channels are even worth the expensive test. Killing attribution because it isn't causal would be throwing out the map because it isn't the territory. Keep the map. Use it to decide where to send the referee. Then let the holdout, not the report, tell you what's true.
Key Takeaways
- ✓The closer a channel sits to the purchase, the more credit it gets and the less it actually caused
- ✓High-ROAS proximity channels can't scale because they harvest a fixed pool of already-decided buyers
- ✓Attribution measures involvement, not causation; only a randomized holdout separates the two
- ✓MTA redistributes credit across observed touches using correlational rules; it has no counterfactual
- ✓Proximity channels are often worth running at a modest budget and useless to scale as acquisition
If retargeting and branded search aren't causal, should I just turn them off?▼
How is this different from multi-touch attribution? Doesn't MTA already solve the credit problem?▼
What's the cheapest holdout I can run this quarter?▼
My best channel has a 6x ROAS. Why would scaling it make ROAS worse rather than just flat?▼
How big does my business need to be to measure incrementality reliably?▼
Can a late touchpoint ever be genuinely causal?▼
How mature is your marketing measurement?
The free Measurement Maturity Assessment shows where you stand, where you're exposed, and what to fix first. 10 questions, 3 minutes.
Take the AssessmentReady to try server-side attribution?
Set up in 10 minutes. Free up to 30K records/month.