Marketing incrementality testing: find out what the campaign changed
Quick answer
Would those customers have bought anyway? That’s the question incrementality testing tries to answer. Compare what happens with a campaign against what happens without it, using groups that give you a fair comparison.
Attribution answers a different question
If someone sees an ad and later buys, an attribution system may assign credit to the ad. The customer might also have bought without seeing it. Incrementality asks about that missing alternative.
This distinction matters most where marketing reaches people already close to purchasing. Retargeting, branded search and customer campaigns can report attractive conversion rates because their audiences start with strong intent. That does not mean the activity has no value; it means the value needs a comparison.
Google’s Meridian GeoX documentation describes geographic experiments as a way to estimate incremental effects and inform marketing mix models. A geographic test is one option, not a substitute for thinking through the business conditions that could confound the result.
Define the decision before the design
Write down what the test should change: whether to maintain a channel, expand a campaign, revise frequency or shift a budget. Pick one primary outcome that reflects that decision. Revenue or contribution may be appropriate; qualified opportunities may be an earlier measure when sales cycles are long.
Set the unit of assignment. It might be a person, account or region, depending on what can be controlled and measured. Avoid assigning individuals independently if colleagues in the same buying group share exposure and outcomes; that can contaminate the comparison.
Specify the treatment and control experience in operational terms. “More marketing” is not a treatment. A defined media budget, audience rule and creative set is something another team can reproduce.
| Design choice | Specify | Failure to avoid |
|---|---|---|
| Outcome | One primary business measure | Choosing the best-looking metric later |
| Assignment | Person, account or geography | Contamination between groups |
| Duration | Observation and stopping rule | Stopping at the first favourable result |
| Analysis | Uncertainty and exclusions | Treating the point estimate as certain |
Protect the comparison
Random assignment helps balance observed and unobserved differences, but implementation still matters. Check whether the treatment actually reached its assigned audience, whether the control was exposed elsewhere and whether sales effort changed unevenly between groups.
For geographic experiments, examine historical similarity, other campaigns, distribution changes, local events and spillover between markets. A region with a new store opening is not a clean comparison for one without it.
Choose the analysis window and stopping rule before seeing the result. Repeatedly checking and stopping when a favourable number appears can exaggerate apparent effects. Plan sample size or detectable effect with appropriate analytical support, especially when the outcome is rare.
Calculate lift and show its uncertainty
Take a simple test with 10,000 people randomly assigned to see the campaign and another 10,000 in a control group. If 240 in the first group buy and 200 in the second do, the conversion rates are 2.4% and 2.0%. That’s a difference of 0.4 percentage points, or a 20% lift relative to the control.
At the treatment group’s size, that corresponds to an estimated 40 additional purchases in this simple comparison. If the campaign cost £4,000, the estimated cost per incremental purchase is £100. It would be wrong to divide by all 240 purchases and call that incremental acquisition cost.
The calculation is not a statistical conclusion. Report confidence or credible intervals under the chosen method, examine assignment and missing data, and state whether the evidence is strong enough for the decision. An uncertain result can still rule out an implausibly large claim without proving the effect is zero.
An uncomfortable result can be a useful one
If a test finds less lift than the platform claims, don’t rush to blame the test. It may be showing that many customers would have bought anyway. Or the effect may be too small for the test to measure reliably. Either way, you’ve learned something the attribution report couldn’t tell you.
Use the result to adjust a specific decision. A profitable average return at the current budget does not establish that the next pound will perform equally well. Google’s response-curve guidance distinguishes average from marginal return and discusses extrapolation risk.
Keep the experiment record alongside the data-science evaluation plan. Preserve the design, dates, exclusions and limitations so future analysts understand what was tested. For smaller teams, a feasible, well-executed experiment is more valuable than an elaborate model supported by weak data. Our guide to choosing marketing evidence helps frame that trade-off.
Related reading
Display advertising quality: what to check beyond cheap CPMs · Marketing attribution vs MMM vs incrementality: which question are you asking? · Marketing budget allocation: fund the next useful decision
Photo: Aaron Lefler / Unsplash.
