Homemade Pasta · Research
Two studies asked what branded search is worth. eBay switched it off and 99.5% of the traffic returned through organic search for free. Edmunds randomized the same shutoff across US markets and lost more than half. Neither is wrong. The gap between them is the finding, and the best available explanation has more to do with the advertiser than the channel.
Why this exists
Homemade Pasta runs paid media for clients in categories with little in common. All of them eventually ask the same thing: how much of this would have happened anyway. The answers on offer are a platform dashboard, which has an interest in the answer, or an industry benchmark, which was measured on somebody else.
We went looking for the benchmark and came back without one. The literature offers something more useful instead. A run of large field experiments, from 2012 to this year, disagree with each other sharply enough, and for reasons interpretable enough, that the disagreement carries the information.
This page reads that evidence, explains what the variance is actually measuring, and gives you the arithmetic to work out whether you can measure it yourself. It crosses from media measurement into marketing partway through, because that is where the evidence goes.
Executive summary
Three findings determine how far your reported numbers can be trusted, why better attribution has not fixed them, and what the variance is actually telling you.
Switch off paid search on your own brand name and organic picks up the traffic, or it does not. eBay kept almost all of it. Edmunds lost more than half. Both measured the same thing the same way.
Travel sits at the bottom of the band, e-commerce at the top. One advertiser can bank three quarters of a last-click conversion. The other has to write off more than half.
The practical consequence is that incrementality works as a diagnostic and fails as a scorecard. A low number rarely means the media was bought badly. It usually means the demand already existed and something else created it, which is a marketing problem. Moving budget between platforms will not touch it. Changing what the advertising is asked to do might.
Measuring it yourself is the only route to a number that applies to you, and it is harder than vendor material suggests. Across 25 large field experiments the median confidence interval on return ran more than 100 percentage points wide. Part Three gives the arithmetic for working out, before you commit delivery, whether your own test can resolve anything at all.
Recommendations
| Action | Why it matters | What it changes |
|---|---|---|
| 1. Stop benchmarking against anyone else's number | Every published figure was measured on a specific brand, category and funnel position. The two cleanest experiments on branded search disagree by a factor of two, and the second set of authors wrote in print that they do not know why. | Removes a class of decision made on a number that was never about you. |
| 2. Run the power arithmetic before you run the test | Most incrementality tests are decided before they start. The effect being hunted is often smaller than the smallest effect the test could detect, so the result is a wide interval around zero and a false conclusion that the channel does nothing. | Roughly 19,600 conversions in the window are needed to resolve a 5% lift, at a 20% holdout, 95% confidence and 80% power. Below about 2,000, only very large effects are visible. |
| 3. Test at the level where the volume is | Incrementality is a portfolio and channel question. Campaign-level tests almost never carry enough conversions, and splitting a fixed volume into more cells lowers power on all of them. | Converts an unanswerable test into an answerable one without spending more. |
| 4. Read your results by new, dormant and existing buyers | The best-supported account of the spread is who the advertising reaches. Effects concentrate in people with no purchase history and people long dormant, and fall close to zero among frequent buyers. This is reading one test three ways after the fact, not running three tests, which is why it does not cost power the way extra cells do. The subgroup estimates are still noisier than the whole. | Shows which share of spend is creating demand and which is collecting it. |
| 5. Decide what share of budget is meant to create demand | If the diagnosis is that advertising is collecting demand somebody else created, no amount of media optimization moves the number. The decision is whether creating demand is a job the marketing is funded to do, and who owns it. That decision sits above the media plan. | Turns a measurement finding into an owned objective. Without it, a low number gets re-litigated every quarter and answered with channel changes that cannot fix it. |
| 6. Treat platform lift tools as measurement, not reporting | Google states that Conversion Lift intentionally ignores your attribution settings. Attributed conversions and incremental conversions are different products from the same vendor, and only one of them answers this question. | Google reduced the stated minimum spend for an incrementality test from upwards of $100,000 to $5,000 in November 2025. The binding constraint has moved from cost to conversion volume. |
Part One
Two shutoffs that disagree, a spread inside a single platform, and what happened when the strongest available models were pointed at the problem.
1.1 The disagreement
In 2012 eBay switched off paid search on its own brand terms and watched what happened to traffic. Edmunds repeated the exercise three years later. The designs differ and it matters: eBay suspended brand advertising on two search engines and used its continued spend on a third as the control, while Edmunds randomized 105 of 210 US media markets. Both measured the same quantity, the share of forgone paid clicks that arrived through organic search instead.
The eBay result is the more famous one, and it carries more weight: it is peer reviewed, the Edmunds paper is still a working draft eight years on. Neither result has been shown to be wrong, and that is the point. It appeared in Econometrica, and its non-brand arm produced the single cleanest demonstration of the problem in the literature: measured experimentally, the return on non-brand search was −63%, with a 95% interval of −124% to −3%. The same data analyzed the way a dashboard analyzes it, with market and date controls, returned +1,632%. Without controls, +4,173%. The two methods do not disagree about how large the return was. They disagree about whether the money worked at all.
The Edmunds authors knew they were contradicting a famous result and said so directly. Their stated conclusion was that money spent on search marketing may be more effective than previously documented, and that they did not know why their answer differed from eBay's. That sentence has not been resolved in the decade since.
The disagreement is not confined to search, or to comparisons across a decade. The largest study of its kind matched 2,226 randomized advertising experiments on Meta against what seven-day last-click attribution reported for the same campaigns. Across everything, about 69% of last-click conversions per dollar were incremental. Split by vertical, the number moves a long way.
One number deserves more attention than the headline. Seven-day last-click explained 19% of the out-of-sample variance in true incrementality overall. Within e-commerce it explained 90%. Within retail and travel the explanatory power was negative, meaning the metric performed worse than simply guessing the average. The same field on the same dashboard is a good guide for one advertiser and an actively misleading one for another.
A reading that only found overstatement would be selective. Two results point the other way, and both need their conditions attached.
A measurement vendor published 640 incrementality tests on Meta and reported that for every $100 of attributed direct-to-consumer revenue, $115 was incremental, an understatement rather than an overstatement.14 Treat it as a pointer, with the caveats attached. The firm sells incrementality testing, no confidence intervals or market-pair criteria are disclosed, and the comparison is between a holdout and a seven-day click window, which measure different things. A holdout captures total business impact including halo and untracked paths. A click window captures the tracked click path. The difference between them is not attribution inflation, and reading it as such is the error in either direction.
In mobile apps the spillover runs positive. When a large game developer shut off install advertising globally, organic installs fell 20% to 30%, because paid volume was feeding app store category rankings that then produced organic discovery.10 The authors conclude that developers in that market may be systematically underinvesting.
Both the direction of the error and its size move with the case.
1.2 The modeling attempt
The obvious response to unreliable attribution is better attribution. That response has been tested at a scale no advertiser could fund, and it failed.
Researchers took 563 randomized experiments run on Meta, producing 663 treatment and control pairs and 1,673 measured outcomes. They then attempted to recover the experimental answer using only observational data, with more than 5,000 user-level features per person and double machine learning. That is a richer feature set than any advertiser or measurement vendor has access to, and a stronger method than any dashboard uses.
The authors' own summary is unusually direct: despite access to large-scale experiments and rich user-level data, they were unable to reliably estimate an advertising campaign's causal effect. The median absolute gap between the model estimate and the experimental one ran 115, 103 and 57 percentage points from the top of the funnel to the bottom.
Read those gaps carefully, because they run in opposite directions depending on the frame. In absolute terms the gap narrows down the funnel, from 115 points to 57. Measured against the effect being estimated it widens: the model returns roughly three times the experimental figure at the top of the funnel and almost five times it at the bottom. Attribution is least trustworthy exactly where budgets are decided.
The more useful finding is what happened when the method improved. This was a much harder test than an earlier study by the same group, which had 15 experiments and simple matching methods and found that half its estimates were off by a factor of three.3 The better method with the richer data did not do better. That pattern, a failure that survives improvement, is the closest thing to a settled result in this literature.
Not everything an ad platform calls a test produces a clean comparison. A 2025 study examined 3,204 lift tests and 181,890 A/B tests on Meta, covering tens of billions of user observations.
The lift tests held up. Only 0.16% of measured imbalances between treatment and control exceeded the conventional threshold, and the distribution behaved as randomization requires.
The A/B tests failed the same check. Across the measured comparisons, 25% of test statistics were significant at the 5% level before any treatment effect, and 22% of imbalance measures exceeded the conventional threshold. The cause is divergent delivery: the platform optimizes each arm toward different people, so the arms stop being comparable. Restricting to awareness objectives with identical configuration cuts the imbalance without clearing it, and the authors state that no configuration eliminates divergent delivery entirely.7
The practical rule is that a lift test and an in-platform A/B test are different instruments. Only the first is measuring incrementality. The evidence here covers Meta, and no equivalent public audit exists for the other platforms.
Part Two
The spread tracks who the advertising reaches, and whether that advertising is making demand or meeting it.
2.1 The mechanism
The eBay study contains a finding that gets quoted far less than its headline, and it is the one that explains the spread.
The effect of paid search was concentrated in two groups: users with no prior purchase, and users who had been dormant for over a year. Among frequent buyers it was close to zero. And frequent buyers, in the authors' own words, account for most paid search traffic and most attributed sales.
So eBay's advertising was working. It was working on a small minority of the people it reached, and being credited for a large majority who were arriving regardless. Average measured incrementality came out near zero not because the channel was ineffective but because the denominator was full of demand that already existed.
Strong brand, high organic presence, buyers already in the consideration set. Advertising intercepts a journey that was happening. The number is low, and moving budget between platforms will not raise it.
Unknown brand, thin organic presence, buyers who were not in market. Advertising starts the journey. The number is high, and the constraint is usually reach and message, with measurement the easy part.
The eBay authors state the boundary of their own finding plainly. They suspect the result generalizes to well known brands already in most consumers' consideration sets, and write that it may not be true for small and new entities with no brand recognition. That is the same axis, named by the people who found it.
It also fits the direction reversal. In mobile apps, where shutting off paid installs cut organic installs by 20% to 30%, advertising was manufacturing the discovery mechanism rather than intercepting it. The funnel gradient is consistent with it too, though more weakly: median measured lift of 29% at the top against 5% at the bottom is what creating demand versus collecting it would look like inside one campaign, and it is also what differing base rates and a longer causal chain would produce. That one does not discriminate between the explanations.
This reading is an interpretation rather than a measured finding, and it should be treated that way. No published study has tested brand strength as the explanation for why eBay and Edmunds disagree, and the Edmunds authors declined to offer one. Other accounts survive the evidence too. Competitors bidding on a brand term change where lost traffic goes and have nothing to do with brand strength. The results page itself was rebuilt between 2012 and 2015. A marketplace and a lead-generation site do not mean the same thing by a recaptured visit. What can be said is that the heterogeneity result is real, and that it is the most direct account on offer.
A low incrementality number is usually read as a media failure and answered with a media response: change the platform, change the targeting, cut the budget. If the mechanism above is right, that response addresses the wrong layer. Advertising that collects existing demand is doing what it was set up to do. The number is low because something else in the business created the demand first, and the advertising is being asked to take credit for it.
The question that follows is not which channel to cut. It is how much of the marketing is expected to create demand at all, who is responsible for that, and whether any of it is currently being measured.
Part Three
Whether your volume can resolve anything, and the practices that survive the evidence.
3.1 Before you test
The case for running your own test is strong. The case for checking whether it can work first is stronger, and it is routinely skipped.
Across 25 large advertising field experiments covering millions of customers, the median confidence interval on return on investment was more than 100 percentage points wide. Individual purchase behavior is volatile enough that separating a profitable campaign from a worthless one can require more than ten million person-weeks of exposure. Underpowered tests come back looking like results. A wide interval around zero reads to most audiences as proof that the channel does nothing.
The smallest lift a holdout test could reliably detect, given the conversions you have to work with.
This test resolves a lift of about 7.8% or more. For scale only, the median lower-funnel lift across 563 randomized Meta experiments was 5%, so an effect of that size would come back inconclusive here.
The arithmetic is unforgiving in a specific way. Precision improves with the square root of volume, so halving the smallest detectable effect requires roughly four times the conversions. An advertiser with 2,000 conversions in the window cannot reach the same answer as one with 20,000 by being more careful. They need a different question, a longer window, or a larger unit of analysis.
That is the practical argument for testing the channel or the whole portfolio rather than individual campaigns. Splitting a fixed conversion volume across more cells lowers power on every one of them, and a set of inconclusive campaign tests is worse than one conclusive channel test, because it looks like evidence.
Some programs cannot get there. If the calculator returns a figure several times the size of any effect anyone has measured, and a longer window or a wider unit of analysis will not close the distance, the honest answer is not to run the test. Spend the holdout on delivery instead, reason from the structure of the business using Part Two, and revisit when volume supports it. A test that cannot resolve the effect still produces a number, and that number will look like a zero somebody can act on.
The argument against borrowing somebody else's figure applies, with a delay, to your own. If incrementality tracks how much demand the advertising creates rather than collects, then it moves when that relationship moves: as the brand becomes better known, as organic presence grows, as a competitor starts bidding on your terms, as the category's demand cycle turns. A result from two years ago was measured on a different business.
That sets a cadence. Re-measure when something structural changes, not on a calendar, and treat the direction of travel between two tests as more informative than either point estimate. A measured number is worth more than a borrowed one because you know the conditions it came from.
The difficulty sits in the ratio of the effect to the noise around it.
In the 25-experiment study above, individual sales had a coefficient of variation around 10, meaning the standard deviation of spend per person was roughly ten times the mean. For a campaign delivering a healthy 25% return, the effect being detected was about $0.35 per person against a mean of $7 and a standard deviation of $75.
Expressed as model fit, a highly profitable campaign in that setting corresponds to an R² on the order of 0.0000054. Any observational method claiming to recover the effect would have to be correctly specified to a tolerance far beyond that, in a setting where selection effects can run thirty times larger than the treatment effect itself. The authors call this an impossible statistical feat, and it is the same wall the modeling study in Part One ran into from the other direction.6
Even their best-powered single experiment, with a 3.5 million person control group and pre-period covariates, still returned a 95% interval on return 60 percentage points wide.
3.2 Practices
The practices the evidence backs, the ones where the trap sits in the reading rather than the doing, and the ones that do not produce an incrementality number at all.
Supported by the evidence above.
It determines whether the test is worth running at all, and it takes minutes. The alternative is spending a quarter of delivery to produce a wide interval around zero.
Both assign exposure rather than optimizing toward it. A 2025 audit of 3,204 Meta lift tests found the randomization held, with imbalance above the conventional threshold in only 0.16% of cases. No equivalent public audit exists for other platforms or for geo designs.
Conversion volume is the binding constraint. Every additional cell divides the same volume and lowers power across all of them.
A test shorter than the consideration period measures the fast converters only, which biases toward the people who were coming anyway and understates the effect.
This is where the eBay effect lived, and it is the difference between a channel that does nothing and a channel that does something for a minority of the people it reaches. Read one test three ways after the fact rather than running three tests. Splitting the design costs power. Splitting the readout does not, though the subgroup estimates are noisier than the whole.
Nothing wrong with the practice. The trap is in what gets concluded from it.
Both are worth having. A geo holdout captures total business impact including halo and untracked paths, a seven-day click window captures the tracked click path. Reading the difference between them as attribution inflation is the error, in either direction.
With a median return interval over 100 percentage points wide across 25 large experiments, an inconclusive result is the expected outcome of a normal-sized test, not a finding about the channel.
A test measures one channel at one moment, under the brand strength, competitive set and demand conditions of that moment. The estimate ages. Re-measure when something structural changes, not on a calendar, and read the movement between two tests instead of defending either one.
These produce a number. None of them produce an incrementality number.
The two cleanest tests on branded search disagree by a factor of two, and within a single platform the incremental share of last-click conversions runs from 48% to 76% by vertical.
Attempted at scale with 563 experiments, 5,000 features per user and double machine learning. The median gap to the experimental answer ran 115, 103 and 57 percentage points by funnel stage.
In the same 2025 audit, a quarter of the measured comparisons across 181,890 in-platform A/B tests were significant at the 5% level before any treatment effect, because delivery optimizes each arm toward different people.
The window governs which conversions get counted, not which ones would have happened anyway. Narrowing it moves the reported number without changing the underlying quantity.
Appendix
Measurement notes, the limits of the evidence, and every source behind the figures above.
A.1 Measurement notes
Google states the distinction in its own documentation more plainly than most vendors would. Standard reporting counts conversions according to the tracking settings and attribution rules configured for each conversion action. Conversion Lift, in Google's words, intentionally ignores those settings and counts separately the net new conversions driven by ad exposure.11
Both numbers appear in the same interface. Only one of them answers the question in this document, and the other one is the default.
Neither Google nor Meta publishes a typical lift magnitude. Google's only quantitative claim in this territory is a 2011 finding that 89% of paid search clicks were incremental, drawn from more than 400 studies of paused accounts.13 It measures clicks rather than conversions, it counts only recapture through organic search, the accounts were paused for their own reasons rather than randomly, and it sits in direct tension with the branded search experiment in Part One.
What has changed recently is access. In November 2025 Google reduced the stated minimum spend for an incrementality test from upwards of $100,000 to $5,000 and added test sizing controls. It also claims conclusive results up to 50% more frequently, an unquantified vendor figure that should be read as one.12 The change removes cost as the barrier to entry. It does not change the conversion volume a conclusive test requires, which is the subject of the calculator above.
Platforms have also begun conceding the gap in their own research. A 2026 paper from TikTok's measurement team states that paid-attributed conversions may systematically overstate true incremental growth where paid channels overlap with organic demand or brand-driven traffic.8 An Amazon white paper describes maintaining a database of hundreds of thousands of randomized tests to calibrate its attribution, while noting that randomized tests alone are too imprecise and too coarse to attribute individual touchpoints.9
The two largest studies here, covering 563 and 2,226 randomized experiments, both draw on the same platform and the same window, November 2019 to March 2020. That period precedes app tracking restrictions, automated campaign types and modeled conversions. Both the attribution side and the delivery side of the comparison have been rebuilt since, and neither paper addresses what that does to its estimates.
The 2,226-experiment study reports a central ratio and vertical breakdowns but no distribution: no percentiles, no variance, no share of campaigns measuring zero or negative lift. A point estimate without a published spread cannot responsibly become a planning multiplier, which is part of why this document does not offer one.
Advertisers who run lift tests are large, sophisticated and self-selected. The authors of one study warn explicitly against generalizing their results to advertisers who did not experiment.
Most of the evidence here was produced by the platforms being measured. Sources 4, 5 and 7 are co-authored with Meta researchers on Meta data, source 9 is an Amazon paper, and sources 11 to 13 are Google. The standard applied to the vendor in source 14 applies to them too. What makes the pattern credible is not any single source but that the failures replicate across independent research groups and across competing platforms.
Finally, the explanation offered in Part Two is an interpretation. The heterogeneity finding it rests on is measured. The extension of it to explain why two branded search experiments disagree is not, and the authors of the second experiment declined to offer any explanation at all.
A.2 Sources
Three numbers that circulate in this area could not be traced to a primary source and are not used anywhere above: an error range of 488% to 948% attributed to the 2023 study in a third-party white paper but absent from the paper itself; a 25% average calibration difference credited to an unnamed vendor white paper; and a 31% efficiency claim for a geo testing framework that appears only in secondary coverage.