Homemade Pasta · Research

Creating demand, or collecting it

Two studies asked what branded search is worth. eBay switched it off and 99.5% of the traffic returned through organic search for free. Edmunds randomized the same shutoff across US markets and lost more than half. Neither is wrong. The gap between them is the finding, and the best available explanation has more to do with the advertiser than the channel.

PublishedSeptember 2026
ReadAbout 15 minutes

Why this exists

Every client asks the same question. The honest answer is not a number.

Homemade Pasta runs paid media for clients in categories with little in common. All of them eventually ask the same thing: how much of this would have happened anyway. The answers on offer are a platform dashboard, which has an interest in the answer, or an industry benchmark, which was measured on somebody else.

We went looking for the benchmark and came back without one. The literature offers something more useful instead. A run of large field experiments, from 2012 to this year, disagree with each other sharply enough, and for reasons interpretable enough, that the disagreement carries the information.

This page reads that evidence, explains what the variance is actually measuring, and gives you the arithmetic to work out whether you can measure it yourself. It crosses from media measurement into marketing partway through, because that is where the evidence goes.

Executive summary

There is no benchmark. There is a reason there is no benchmark.

Three findings determine how far your reported numbers can be trusted, why better attribution has not fixed them, and what the variance is actually telling you.

Brand clicks won back by organicSame shutoff, two advertisers
eBay, 201299.5%
Edmunds, 2015under 50%

Switch off paid search on your own brand name and organic picks up the traffic, or it does not. eBay kept almost all of it. Edmunds lost more than half. Both measured the same thing the same way.

  1. No external incrementality figure survives contact with a different advertiser. When eBay switched off branded search, 99.5% of the forgone clicks arrived through organic search instead. When Edmunds ran a randomized version of the same shutoff, more than half the traffic went elsewhere. Inside Meta alone, over a single five-month window, the share of last-click conversions that proved incremental ran from 76% in e-commerce to 48% in travel.
  2. Better modeling has been tried properly and did not recover the experimental answer. One study took 563 randomized experiments and tried to reproduce their results from observational data alone, using more than 5,000 features per user and double machine learning. Median measured lift was 29%, 18% and 5% moving down the funnel. The models returned 83%, 58% and 24%. The authors concluded they were unable to reliably estimate a campaign's causal effect.
  3. The variance says more about the advertiser than about the method. eBay's paid search moved people with no prior purchase and people dormant for over a year. It did close to nothing for frequent buyers, who were most of the traffic and most of the attributed sales. That is the difference between advertising that creates demand and advertising that collects it. The heterogeneity is measured. Reading it as the explanation for the whole spread is our inference, and Part Two says so plainly.
Last-click incrementality rangeMeta, across verticals, one five-month window
48% 76% 0%25%50%75%100%

Travel sits at the bottom of the band, e-commerce at the top. One advertiser can bank three quarters of a last-click conversion. The other has to write off more than half.

The practical consequence is that incrementality works as a diagnostic and fails as a scorecard. A low number rarely means the media was bought badly. It usually means the demand already existed and something else created it, which is a marketing problem. Moving budget between platforms will not touch it. Changing what the advertising is asked to do might.

Measuring it yourself is the only route to a number that applies to you, and it is harder than vendor material suggests. Across 25 large field experiments the median confidence interval on return ran more than 100 percentage points wide. Part Three gives the arithmetic for working out, before you commit delivery, whether your own test can resolve anything at all.

Recommendations

Mise en place, in order of return

ActionWhy it mattersWhat it changes
1. Stop benchmarking against anyone else's number Every published figure was measured on a specific brand, category and funnel position. The two cleanest experiments on branded search disagree by a factor of two, and the second set of authors wrote in print that they do not know why. Removes a class of decision made on a number that was never about you.
2. Run the power arithmetic before you run the test Most incrementality tests are decided before they start. The effect being hunted is often smaller than the smallest effect the test could detect, so the result is a wide interval around zero and a false conclusion that the channel does nothing. Roughly 19,600 conversions in the window are needed to resolve a 5% lift, at a 20% holdout, 95% confidence and 80% power. Below about 2,000, only very large effects are visible.
3. Test at the level where the volume is Incrementality is a portfolio and channel question. Campaign-level tests almost never carry enough conversions, and splitting a fixed volume into more cells lowers power on all of them. Converts an unanswerable test into an answerable one without spending more.
4. Read your results by new, dormant and existing buyers The best-supported account of the spread is who the advertising reaches. Effects concentrate in people with no purchase history and people long dormant, and fall close to zero among frequent buyers. This is reading one test three ways after the fact, not running three tests, which is why it does not cost power the way extra cells do. The subgroup estimates are still noisier than the whole. Shows which share of spend is creating demand and which is collecting it.
5. Decide what share of budget is meant to create demand If the diagnosis is that advertising is collecting demand somebody else created, no amount of media optimization moves the number. The decision is whether creating demand is a job the marketing is funded to do, and who owns it. That decision sits above the media plan. Turns a measurement finding into an owned objective. Without it, a low number gets re-litigated every quarter and answered with channel changes that cannot fix it.
6. Treat platform lift tools as measurement, not reporting Google states that Conversion Lift intentionally ignores your attribution settings. Attributed conversions and incremental conversions are different products from the same vendor, and only one of them answers this question. Google reduced the stated minimum spend for an incrementality test from upwards of $100,000 to $5,000 in November 2025. The binding constraint has moved from cost to conversion volume.

Part One

Nobody else's number is yours

Two shutoffs that disagree, a spread inside a single platform, and what happened when the strongest available models were pointed at the problem.

1.1 The disagreement

Same question, same channel, opposite answers

In 2012 eBay switched off paid search on its own brand terms and watched what happened to traffic. Edmunds repeated the exercise three years later. The designs differ and it matters: eBay suspended brand advertising on two search engines and used its continued spend on a third as the control, while Edmunds randomized 105 of 210 US media markets. Both measured the same quantity, the share of forgone paid clicks that arrived through organic search instead.

Point estimate Reported range
eBay 2012, Econometrica 99.5% Edmunds 2015, working paper under 50% 0% 25% 50% 75% 100% Share of forgone paid clicks recaptured through organic search
eBay recaptured almost everything, so the paid clicks were close to worthless on brand terms. Edmunds lost more than half, and in the markets where its brand was already best known it recaptured only about 28%. The band shown spans those two reported figures.1, 2

The eBay result is the more famous one, and it carries more weight: it is peer reviewed, the Edmunds paper is still a working draft eight years on. Neither result has been shown to be wrong, and that is the point. It appeared in Econometrica, and its non-brand arm produced the single cleanest demonstration of the problem in the literature: measured experimentally, the return on non-brand search was −63%, with a 95% interval of −124% to −3%. The same data analyzed the way a dashboard analyzes it, with market and date controls, returned +1,632%. Without controls, +4,173%. The two methods do not disagree about how large the return was. They disagree about whether the money worked at all.

The Edmunds authors knew they were contradicting a famous result and said so directly. Their stated conclusion was that money spent on search marketing may be more effective than previously documented, and that they did not know why their answer differed from eBay's. That sentence has not been resolved in the decade since.

One platform, three answers

The disagreement is not confined to search, or to comparisons across a decade. The largest study of its kind matched 2,226 randomized advertising experiments on Meta against what seven-day last-click attribution reported for the same campaigns. Across everything, about 69% of last-click conversions per dollar were incremental. Split by vertical, the number moves a long way.

All advertisers 69% E-commerce 76% Retail 63% Travel 48% 0% 25% 50% 75% 100% Share of last-click conversions per dollar that were incremental
Across 2,226 randomized experiments on Meta. In travel, roughly half of last-click conversions would have happened without the ad. The same attribution metric predicted true incrementality well in e-commerce and worse than useless in retail and travel.5

One number deserves more attention than the headline. Seven-day last-click explained 19% of the out-of-sample variance in true incrementality overall. Within e-commerce it explained 90%. Within retail and travel the explanatory power was negative, meaning the metric performed worse than simply guessing the average. The same field on the same dashboard is a good guide for one advertiser and an actively misleading one for another.

Where the direction reverses

A reading that only found overstatement would be selective. Two results point the other way, and both need their conditions attached.

A measurement vendor published 640 incrementality tests on Meta and reported that for every $100 of attributed direct-to-consumer revenue, $115 was incremental, an understatement rather than an overstatement.14 Treat it as a pointer, with the caveats attached. The firm sells incrementality testing, no confidence intervals or market-pair criteria are disclosed, and the comparison is between a holdout and a seven-day click window, which measure different things. A holdout captures total business impact including halo and untracked paths. A click window captures the tracked click path. The difference between them is not attribution inflation, and reading it as such is the error in either direction.

In mobile apps the spillover runs positive. When a large game developer shut off install advertising globally, organic installs fell 20% to 30%, because paid volume was feeding app store category rankings that then produced organic discovery.10 The authors conclude that developers in that market may be systematically underinvesting.

Both the direction of the error and its size move with the case.

1.2 The modeling attempt

Five thousand signals, and still the wrong answer

The obvious response to unreliable attribution is better attribution. That response has been tested at a scale no advertiser could fund, and it failed.

Researchers took 563 randomized experiments run on Meta, producing 663 treatment and control pairs and 1,673 measured outcomes. They then attempted to recover the experimental answer using only observational data, with more than 5,000 user-level features per person and double machine learning. That is a richer feature set than any advertiser or measurement vendor has access to, and a stronger method than any dashboard uses.

Randomized experiment Model estimate on the same campaigns
Upper funnel 29% 83% Mid funnel 18% 58% Lower funnel 5% 24% 0% 30% 60% 90% Median measured lift, by funnel stage
Lower-funnel outcomes are purchases and checkouts. Upper-funnel outcomes are registrations and page views. The model overstates at every stage. Relative to the effect being measured, it overstates worst at the bottom.4

The authors' own summary is unusually direct: despite access to large-scale experiments and rich user-level data, they were unable to reliably estimate an advertising campaign's causal effect. The median absolute gap between the model estimate and the experimental one ran 115, 103 and 57 percentage points from the top of the funnel to the bottom.

Read those gaps carefully, because they run in opposite directions depending on the frame. In absolute terms the gap narrows down the funnel, from 115 points to 57. Measured against the effect being estimated it widens: the model returns roughly three times the experimental figure at the top of the funnel and almost five times it at the bottom. Attribution is least trustworthy exactly where budgets are decided.

The more useful finding is what happened when the method improved. This was a much harder test than an earlier study by the same group, which had 15 experiments and simple matching methods and found that half its estimates were off by a factor of three.3 The better method with the richer data did not do better. That pattern, a failure that survives improvement, is the closest thing to a settled result in this literature.

Choosing between attribution models is the wrong argument. The open question was whether any model could stand in for an experiment, and the largest attempt to date answered it.
Not every test is a test

Not everything an ad platform calls a test produces a clean comparison. A 2025 study examined 3,204 lift tests and 181,890 A/B tests on Meta, covering tens of billions of user observations.

The lift tests held up. Only 0.16% of measured imbalances between treatment and control exceeded the conventional threshold, and the distribution behaved as randomization requires.

The A/B tests failed the same check. Across the measured comparisons, 25% of test statistics were significant at the 5% level before any treatment effect, and 22% of imbalance measures exceeded the conventional threshold. The cause is divergent delivery: the platform optimizes each arm toward different people, so the arms stop being comparable. Restricting to awareness objectives with identical configuration cuts the imbalance without clearing it, and the authors state that no configuration eliminates divergent delivery entirely.7

The practical rule is that a lift test and an in-platform A/B test are different instruments. Only the first is measuring incrementality. The evidence here covers Meta, and no equivalent public audit exists for the other platforms.

Part Two

Demand you made, demand you met

The spread tracks who the advertising reaches, and whether that advertising is making demand or meeting it.

2.1 The mechanism

The ads worked on the people who were not coming anyway

The eBay study contains a finding that gets quoted far less than its headline, and it is the one that explains the spread.

The effect of paid search was concentrated in two groups: users with no prior purchase, and users who had been dormant for over a year. Among frequent buyers it was close to zero. And frequent buyers, in the authors' own words, account for most paid search traffic and most attributed sales.

So eBay's advertising was working. It was working on a small minority of the people it reached, and being credited for a large majority who were arriving regardless. Average measured incrementality came out near zero not because the channel was ineffective but because the denominator was full of demand that already existed.

Collecting demand Low measured lift

Strong brand, high organic presence, buyers already in the consideration set. Advertising intercepts a journey that was happening. The number is low, and moving budget between platforms will not raise it.

Creating demand High measured lift

Unknown brand, thin organic presence, buyers who were not in market. Advertising starts the journey. The number is high, and the constraint is usually reach and message, with measurement the easy part.

The eBay authors state the boundary of their own finding plainly. They suspect the result generalizes to well known brands already in most consumers' consideration sets, and write that it may not be true for small and new entities with no brand recognition. That is the same axis, named by the people who found it.

It also fits the direction reversal. In mobile apps, where shutting off paid installs cut organic installs by 20% to 30%, advertising was manufacturing the discovery mechanism rather than intercepting it. The funnel gradient is consistent with it too, though more weakly: median measured lift of 29% at the top against 5% at the bottom is what creating demand versus collecting it would look like inside one campaign, and it is also what differing base rates and a longer causal chain would produce. That one does not discriminate between the explanations.

This reading is an interpretation rather than a measured finding, and it should be treated that way. No published study has tested brand strength as the explanation for why eBay and Edmunds disagree, and the Edmunds authors declined to offer one. Other accounts survive the evidence too. Competitors bidding on a brand term change where lost traffic goes and have nothing to do with brand strength. The results page itself was rebuilt between 2012 and 2015. A marketplace and a lead-generation site do not mean the same thing by a recaptured visit. What can be said is that the heterogeneity result is real, and that it is the most direct account on offer.

Where this leaves the budget conversation

A low incrementality number is usually read as a media failure and answered with a media response: change the platform, change the targeting, cut the budget. If the mechanism above is right, that response addresses the wrong layer. Advertising that collects existing demand is doing what it was set up to do. The number is low because something else in the business created the demand first, and the advertising is being asked to take credit for it.

The question that follows is not which channel to cut. It is how much of the marketing is expected to create demand at all, who is responsible for that, and whether any of it is currently being measured.

Part Three

The test you can actually run

Whether your volume can resolve anything, and the practices that survive the evidence.

3.1 Before you test

Most incrementality tests are decided before they start

The case for running your own test is strong. The case for checking whether it can work first is stronger, and it is routinely skipped.

Across 25 large advertising field experiments covering millions of customers, the median confidence interval on return on investment was more than 100 percentage points wide. Individual purchase behavior is volatile enough that separating a profitable campaign from a worthless one can require more than ten million person-weeks of exposure. Underpowered tests come back looking like results. A wide interval around zero reads to most audiences as proof that the channel does nothing.

Can your test answer the question?

The smallest lift a holdout test could reliably detect, given the conversions you have to work with.

Scale
8,000
Total across test and control, for the whole test period
20%
Larger holdouts buy precision and cost delivery
Confidence level
Statistical power held at 80% throughout
Smallest lift you could detect 7.8% At 95% confidence and 80% power
Large effects only

This test resolves a lift of about 7.8% or more. For scale only, the median lower-funnel lift across 563 randomized Meta experiments was 5%, so an effect of that size would come back inconclusive here.

Conversions in control1,600
Conversions in test6,400
Needed to resolve a 5% lift19,600
Multiplier applied2.80
Minimum detectable effect, relative: MDE = z × √(1/Ccontrol + 1/Ctest), where Ccontrol and Ctest are the expected conversions in each group, and z is the sum of the critical values for the confidence level and for 80% power (1.96 + 0.84 = 2.80 at 95%). This is the count, or log-ratio, form. It assumes a low conversion rate and is conservative otherwise, and its input is conversions, not user counts. Geo tests assign whole markets rather than people, so outcomes inside a market are correlated and the same conversion volume buys less precision. A geo design will need more conversions than the figure above to resolve the same lift, not fewer.

The arithmetic is unforgiving in a specific way. Precision improves with the square root of volume, so halving the smallest detectable effect requires roughly four times the conversions. An advertiser with 2,000 conversions in the window cannot reach the same answer as one with 20,000 by being more careful. They need a different question, a longer window, or a larger unit of analysis.

That is the practical argument for testing the channel or the whole portfolio rather than individual campaigns. Splitting a fixed conversion volume across more cells lowers power on every one of them, and a set of inconclusive campaign tests is worse than one conclusive channel test, because it looks like evidence.

Some programs cannot get there. If the calculator returns a figure several times the size of any effect anyone has measured, and a longer window or a wider unit of analysis will not close the distance, the honest answer is not to run the test. Spend the holdout on delivery instead, reason from the structure of the business using Part Two, and revisit when volume supports it. A test that cannot resolve the effect still produces a number, and that number will look like a zero somebody can act on.

Your own number ages too

The argument against borrowing somebody else's figure applies, with a delay, to your own. If incrementality tracks how much demand the advertising creates rather than collects, then it moves when that relationship moves: as the brand becomes better known, as organic presence grows, as a competitor starts bidding on your terms, as the category's demand cycle turns. A result from two years ago was measured on a different business.

That sets a cadence. Re-measure when something structural changes, not on a calendar, and treat the direction of travel between two tests as more informative than either point estimate. A measured number is worth more than a borrowed one because you know the conditions it came from.

Why the interval is so wide

The difficulty sits in the ratio of the effect to the noise around it.

In the 25-experiment study above, individual sales had a coefficient of variation around 10, meaning the standard deviation of spend per person was roughly ten times the mean. For a campaign delivering a healthy 25% return, the effect being detected was about $0.35 per person against a mean of $7 and a standard deviation of $75.

Expressed as model fit, a highly profitable campaign in that setting corresponds to an R² on the order of 0.0000054. Any observational method claiming to recover the effect would have to be correctly specified to a tolerance far beyond that, in a setting where selection effects can run thirty times larger than the treatment effect itself. The authors call this an impossible statistical feat, and it is the same wall the modeling study in Part One ran into from the other direction.6

Even their best-powered single experiment, with a 3.5 million person control group and pre-period covariates, still returned a 95% interval on return 60 percentage points wide.

3.2 Practices

Five to run, three to watch, four to drop

The practices the evidence backs, the ones where the trap sits in the reading rather than the doing, and the ones that do not produce an incrementality number at all.

Recommended

Supported by the evidence above.

Run the power calculation before designing the test

It determines whether the test is worth running at all, and it takes minutes. The alternative is spending a quarter of delivery to produce a wide interval around zero.

Use a platform lift test or a geo holdout

Both assign exposure rather than optimizing toward it. A 2025 audit of 3,204 Meta lift tests found the randomization held, with imbalance above the conventional threshold in only 0.16% of cases. No equivalent public audit exists for other platforms or for geo designs.

Test at the channel or portfolio level

Conversion volume is the binding constraint. Every additional cell divides the same volume and lowers power across all of them.

Run the window past the purchase cycle

A test shorter than the consideration period measures the fast converters only, which biases toward the people who were coming anyway and understates the effect.

Split results by new, dormant and existing buyers

This is where the eBay effect lived, and it is the difference between a channel that does nothing and a channel that does something for a minority of the people it reaches. Read one test three ways after the fact rather than running three tests. Splitting the design costs power. Splitting the readout does not, though the subgroup estimates are noisier than the whole.

Risky

Nothing wrong with the practice. The trap is in what gets concluded from it.

Comparing a geo holdout result to a click-window number

Both are worth having. A geo holdout captures total business impact including halo and untracked paths, a seven-day click window captures the tracked click path. Reading the difference between them as attribution inflation is the error, in either direction.

Reading one inconclusive test as proof the channel does nothing

With a median return interval over 100 percentage points wide across 25 large experiments, an inconclusive result is the expected outcome of a normal-sized test, not a finding about the channel.

Treating one test in one season as a structural constant

A test measures one channel at one moment, under the brand strength, competitive set and demand conditions of that moment. The estimate ages. Re-measure when something structural changes, not on a calendar, and read the movement between two tests instead of defending either one.

Avoid

These produce a number. None of them produce an incrementality number.

Adopting a published benchmark as your planning figure

The two cleanest tests on branded search disagree by a factor of two, and within a single platform the incremental share of last-click conversions runs from 48% to 76% by vertical.

Inferring incrementality from a multi-touch attribution model

Attempted at scale with 563 experiments, 5,000 features per user and double machine learning. The median gap to the experimental answer ran 115, 103 and 57 percentage points by funnel stage.

Using in-platform A/B tests as incrementality evidence

In the same 2025 audit, a quarter of the measured comparisons across 181,890 in-platform A/B tests were significant at the 5% level before any treatment effect, because delivery optimizes each arm toward different people.

Changing the attribution window to close an incrementality gap

The window governs which conversions get counted, not which ones would have happened anyway. Narrowing it moves the reported number without changing the underlying quantity.

Appendix

Back of house

Measurement notes, the limits of the evidence, and every source behind the figures above.

A.1 Measurement notes

The other number in the same dashboard

Google states the distinction in its own documentation more plainly than most vendors would. Standard reporting counts conversions according to the tracking settings and attribution rules configured for each conversion action. Conversion Lift, in Google's words, intentionally ignores those settings and counts separately the net new conversions driven by ad exposure.11

Both numbers appear in the same interface. Only one of them answers the question in this document, and the other one is the default.

The platforms’ own disclosures

Neither Google nor Meta publishes a typical lift magnitude. Google's only quantitative claim in this territory is a 2011 finding that 89% of paid search clicks were incremental, drawn from more than 400 studies of paused accounts.13 It measures clicks rather than conversions, it counts only recapture through organic search, the accounts were paused for their own reasons rather than randomly, and it sits in direct tension with the branded search experiment in Part One.

What has changed recently is access. In November 2025 Google reduced the stated minimum spend for an incrementality test from upwards of $100,000 to $5,000 and added test sizing controls. It also claims conclusive results up to 50% more frequently, an unquantified vendor figure that should be read as one.12 The change removes cost as the barrier to entry. It does not change the conversion volume a conclusive test requires, which is the subject of the calculator above.

Platforms have also begun conceding the gap in their own research. A 2026 paper from TikTok's measurement team states that paid-attributed conversions may systematically overstate true incremental growth where paid channels overlap with organic demand or brand-driven traffic.8 An Amazon white paper describes maintaining a database of hundreds of thousands of randomized tests to calibrate its attribution, while noting that randomized tests alone are too imprecise and too coarse to attribute individual touchpoints.9

Where this evidence runs out

The two largest studies here, covering 563 and 2,226 randomized experiments, both draw on the same platform and the same window, November 2019 to March 2020. That period precedes app tracking restrictions, automated campaign types and modeled conversions. Both the attribution side and the delivery side of the comparison have been rebuilt since, and neither paper addresses what that does to its estimates.

The 2,226-experiment study reports a central ratio and vertical breakdowns but no distribution: no percentiles, no variance, no share of campaigns measuring zero or negative lift. A point estimate without a published spread cannot responsibly become a planning multiplier, which is part of why this document does not offer one.

Advertisers who run lift tests are large, sophisticated and self-selected. The authors of one study warn explicitly against generalizing their results to advertisers who did not experiment.

Most of the evidence here was produced by the platforms being measured. Sources 4, 5 and 7 are co-authored with Meta researchers on Meta data, source 9 is an Amazon paper, and sources 11 to 13 are Google. The standard applied to the vendor in source 14 applies to them too. What makes the pattern credible is not any single source but that the failures replicate across independent research groups and across competing platforms.

Finally, the explanation offered in Part Two is an interpretation. The heterogeneity finding it rests on is measured. The extension of it to explain why two branded search experiments disagree is not, and the authors of the second experiment declined to offer any explanation at all.

A.2 Sources

Every ingredient, with its source

Full source list
1Blake, Nosko & Tadelis, “Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment.” Econometrica, January 2015. DOI 10.3982/ECTA12423.Peer reviewedSource of the 99.5% recapture figure, the −63% non-brand return and its interval, the +1,632% and +4,173% observational estimates, and the heterogeneity by purchase history.
2Coviello, Gneezy & Goette, field experiment on paid search effectiveness at Edmunds.com. CEPR Discussion Paper 12333, September 2017, revised October 2018. Still unpublished as of September 2026.Working paperRandomized across 210 US media markets, 105 treated, August to November 2015. Source of the under-50% recapture figure, the 28% figure in high-penetration markets, and the authors' statement that they cannot explain the difference from eBay.
3Gordon, Zettelmeyer, Bhargava & Chapsky, “A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook.” Marketing Science 38(2), 2019, 193–225.Peer reviewed15 randomized experiments on Meta, roughly 500 million user-experiment observations. Source of the finding that half the studies were off by a factor of three. Co-authored with Meta researchers on Meta data.
4Gordon, Moakler & Zettelmeyer, “Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement.” Marketing Science 42(4), 2023, 768–793.Peer reviewed563 experiments, 663 treatment and control pairs, over 5,000 user-level features. Source of the funnel medians, the model estimates and the percentage-point gaps. Co-authored with Meta researchers on Meta data.
5Gordon, Moakler & Zettelmeyer, “Predicted Incrementality by Experimentation (PIE) for Ad Measurement.” NBER Working Paper 35044, April 2026.Working paper2,226 experiment and conversion-event pairs drawn from 839 Meta experiments, November 2019 to March 2020. Source of the 69% incremental share, the vertical breakdowns and the out-of-sample explanatory power figures. Co-authored with Meta researchers on Meta data.
6Lewis & Rao, “The Unfavorable Economics of Measuring the Returns to Advertising.” Quarterly Journal of Economics 130(4), 2015, 1941–1973.Peer reviewed25 large field experiments across 19 retailers and 6 financial services firms. Source of the 100-point confidence interval, the coefficient of variation and the model-fit argument.
7Burtch, Moakler, Gordon, Zhang & Hill, “Characterizing and Minimizing Divergent Delivery in Meta Advertising Experiments.” MSI Working Paper 25-140, November 2025.Working paper3,204 lift tests and 181,890 A/B tests on Meta, begun March to June 2025. Source of the randomization quality comparison between the two instruments. Four of the five authors are at Meta.
8Li, Yuan, Yang, Chen & Song, “Attributed, But Not Incremental.” arXiv:2606.26690, June 2026. Accepted at ADKDD 2026.PreprintAuthored by TikTok's measurement team. Source of the statement that paid-attributed conversions may systematically overstate incremental growth.
9Lewis, Zettelmeyer, Gordon and colleagues, on Amazon Ads multi-touch attribution. arXiv:2508.08209, August 2025.PreprintAmazon white paper with academic co-authors. Source of the calibration approach and the stated limits of randomized tests for touchpoint attribution. Contains no quantitative attributed-versus-incremental comparison on real data.
10Ju, Zhao & Aral, “Advertising Spillovers in Mobile Apps: Evidence from Ad Shutoffs and Store Rankings.” arXiv:2504.16151, April 2025, revised July 2026.PreprintSource of the 20% to 30% fall in organic installs following a global paid shutoff. A single-firm shutoff event study, not a randomized design.
11Google, “Understand your Conversion Lift based on users.” Google Ads Help.Platform documentationSource of the statement that Conversion Lift intentionally ignores standard conversion tracking settings and attribution rules.
12Google, incrementality testing improvements. Google Ads Help, announced 11 November 2025.Platform documentationSource of the reduction in stated minimum test spend from upwards of $100,000 to $5,000, the addition of custom test sizing, and the claim of conclusive results up to 50% more frequently.
13Chan, Yuan, Koehler & Kumar, “Incremental Clicks Impact of Search Advertising.” Google, 2011.Platform researchSource of the 89% incremental clicks figure, drawn from more than 400 studies of paused accounts. Measures clicks rather than conversions, and the accounts were not paused at random.
14Haus, “The Meta Report: Lessons from 640 Haus Incrementality Experiments,” July 2025, and “Is Meta Incremental?”, August 2025.Vendor640 incrementality tests among brands averaging $14 million in annual spend on the platform, averaging 18.6 days of treatment plus an 8.8 day post-treatment window. The $115 per $100 attributed figure appears in the second post. Published by a firm that sells incrementality measurement. No confidence intervals or market-pair criteria disclosed.

Also consulted

·Google Ads Help, “About Conversion Lift.” Definition of the treatment and control design.
·Meta business help documentation on standard and incremental attribution, as quoted by trade press in September 2025. Meta's own pages could not be retrieved directly and the definitions are second-hand.
·Filippou, Quach & Jha, on pseudo-incrementality testing from naturally occurring budget changes. arXiv:2609.18257, September 2026. Method paper, validated on two released benchmark cases.
·Measured, new customer acquisition analysis across 10,000+ campaigns, published June 2026. Reports incremental return figures but no platform-reported comparison, so it does not speak to the gap.

Figures deliberately excluded

Three numbers that circulate in this area could not be traced to a primary source and are not used anywhere above: an error range of 488% to 948% attributed to the 2023 study in a third-party white paper but absent from the paper itself; a 25% average calibration difference credited to an unnamed vendor white paper; and a 31% efficiency claim for a geo testing framework that appears only in secondary coverage.