Guide
Incrementality testing: a complete guide
How to measure what your marketing actually caused, rather than what claimed credit for it. Test designs, statistics without the maths, how results calibrate a marketing mix model, and an honest account of what incrementality testing cannot do.
What incrementality testing is
Incrementality testing measures the true causal effect of marketing by comparing places where you change spend against places where you leave it alone, then reading the difference in business outcomes.
That is the whole idea. You change one thing in one set of regions, keep everything the same in another set, and the gap between them is what your spend caused. Not what a platform claimed, not what a model estimated — what actually happened when the money moved.
The reason this matters is that almost every other measurement method describes correlation and invites you to read it as cause. Retargeting looks excellent because it reaches people already heading for checkout. Branded search looks excellent because it intercepts people typing your name. Both will report strong returns for as long as you keep paying, and neither figure tells you what would have happened if you had stopped.
The problem it solves
Left unchecked, that gap produces three familiar situations:
- Budget stuck in channels that look good and grow nothing. Spend accumulates where measurement is easiest rather than where impact is largest.
- Blackout fear. Nobody will turn anything off, because nobody can say what would break — so nothing is ever tested and the uncertainty compounds.
- Budget meetings with no neutral ground. Marketing cites the platform, finance cites the P&L, and there is no shared evidence to settle it.
An incrementality test produces that shared evidence. It answers questions that dashboards structurally cannot: does this channel work at all, what happens if we spend more, is this new platform worth adding, and can we take money out of branded search without losing sales.
Why it matters now
Attribution lost the benefit of the doubt
Third-party cookies are going, app tracking restrictions have broken user-level measurement, and real customer journeys run across devices and weeks. What survives is a partial record, and last-click and simple multi-touch models treat that partial record as though it were complete. They over-credit channels that are easy to track, under-credit channels that create demand, and present the result with a precision the underlying data cannot support.
MMM came back, and it needs ground truth
Marketing mix modeling answers the top-down question well. It works on aggregate data, so it is privacy-safe and covers channels with no click tracking at all. But it is a model, and models drift. Two different specifications can fit the same history and disagree about what to do next, and nothing inside the model can tell you which is right.
Experiments are what stop that drift. A geo test produces a measured fact — this channel, this market, this spend level, this effect — that the model has to be consistent with. That is why the two methods belong together rather than in competition:
- A unified data model — orders, costs, returns, customers and channels in one place, so outcomes can be read consistently
- Incrementality tests — clean experiments on specific questions
- MMM — a continuous view of the whole mix, calibrated by those tests
- Agents and analysis — day-to-day decisions that run on all three
The move that stack enables is from “we think this channel works” to “we know, because we tested it, and we can see it in the model and in the profit”.
Core concepts
Treatment and control
Every incrementality test compares two groups:
- Control — regions where spend on the tested campaigns stays exactly as it was
- Treatment — regions where spend on those campaigns changes
A country is split into two groups of regions, balanced on historical performance so they behave alike before the test starts. Spend then changes in the treatment group only. If the groups were genuinely comparable at the outset, the difference in outcomes during the test is the causal effect of that change.
The balancing is the part that decides whether a test is worth anything. Two groups that were already diverging will produce a difference regardless of what you do to the spend, and no amount of statistical treatment afterwards will separate the two explanations.
Treatment and control split
One country, divided into two groups of regions matched on historical performance.
Baseline sales, treatment
€1.42m
Baseline sales, control
€1.39m
Groups are interleaved, not north against south. Matching is on historical behaviour, not geography.
Correlation
- What it says
- Two things moved together. “Sales rose when we increased Meta spend.”
- What it cannot tell you
- Whether one caused the other, or something else moved both.
Causation
- What it says
- One thing drove the other, because we changed it deliberately and held everything else steady.
- What it cannot tell you
- How large the effect is in profit terms, unless you measure profit.
Incrementality
- What it says
- The part of the result that exists only because of the marketing — the sales and profit you would have missed.
- What it cannot tell you
- Whether it will still be true next quarter. Results are specific to their setting.
| What it says | What it cannot tell you | |
|---|---|---|
| Correlation | Two things moved together. “Sales rose when we increased Meta spend.” | Whether one caused the other, or something else moved both. |
| Causation | One thing drove the other, because we changed it deliberately and held everything else steady. | How large the effect is in profit terms, unless you measure profit. |
| Incrementality | The part of the result that exists only because of the marketing — the sales and profit you would have missed. | Whether it will still be true next quarter. Results are specific to their setting. |
Statistical significance, without the maths
Not every difference between two groups is real. One group can simply have a lucky fortnight. Significance testing asks a narrow question: is this difference large enough that random variation is an unlikely explanation?
- A significant result means chance is an unlikely explanation. You can act on it.
- A non-significant result is not proof that nothing happened. It means the test could not tell. Treat it as a hypothesis, not a verdict.
That second point is where most testing programmes go wrong. “No significant effect” gets reported as “the channel does not work”, when often it means the test was too short, the spend change too small, or the market too thin to detect an effect that was really there.
Power, and the minimum detectable effect
The question worth asking before a test rather than after is: how large would an effect have to be for this design to see it? That figure is the minimum detectable effect, and working it out is a power analysis. It is the single most useful piece of arithmetic in experiment design and the one most often skipped.
If a market’s volume and your planned spend contrast give a minimum detectable effect of 15%, and the channel realistically drives 6%, the test cannot succeed — it will return “no significant effect” regardless of the truth. Better to widen the spend contrast, extend the window, pick a larger market, or accept that this question is not answerable here. Running it anyway produces a number that looks like a finding and is not one.
Three things drive whether a test can produce an answer at all:
- Time — enough weeks for the effect to appear and settle
- Contrast — a large enough spend difference between the groups
- Volume — enough orders for the difference to rise above noise
If any of the three is missing, the honest answer is to change the design rather than run it and interpret the result generously.
The four test designs that answer most questions
Experiment design can get arbitrarily sophisticated, but nearly every strategic question a commerce brand has fits one of four shapes. Getting the shape right matters more than the statistics — a lift test asks whether a channel works at all, while a marginal test asks what the next euro into it returns, and those have different answers because average return and marginal return are not the same number.
Choosing the design
Start from the question you need answered, not the test you want to run.
Does it work at all?
Lift test
% liftShould we add it?
New channel test
CAC / LTVHow much further?
Marginal spend test
Δ revenueAre we cannibalising?
Branded search test
Cannibalisation %One question per test. Stacking two changes into one window means you cannot attribute the result to either.
Lift test
Does this channel work at all?
- Treatment
- Spend on the tested campaigns is reduced or switched off.
- Control
- Spend continues exactly as before.
- Read it on
- % lift in gross sales and in profit
What to do with it
Run this first on any channel taking meaningful budget.
What it costs you
You give up some revenue in the treatment regions for the duration. That is the price of the answer, and it is the highest of the four designs.
Branded search test
Are we paying for demand we already had?
- Treatment
- Brand bidding is reduced in the treatment regions.
- Control
- Brand bidding continues at normal levels.
- Read it on
- Cannibalisation % — how much paid simply replaced organic
What to do with it
Usually the highest-value first test a brand can run, because the suspected waste is concentrated and the spend is easy to change.
What it costs you
If competitors or resellers bid aggressively on your name, conceding the auction can cost real orders. The test measures that trade-off rather than assuming it away.
New channel test
Should this platform earn a permanent seat?
- Treatment
- The new channel runs in the treatment regions only, alongside the existing mix.
- Control
- The existing mix continues unchanged.
- Read it on
- Incremental CAC and the LTV of the customers it brings
What to do with it
Test one or two new channels a year before committing to scale.
What it costs you
Needs enough budget and platform-appropriate creative to produce a measurable effect. A new channel starved of either will test as ineffective when the design was the problem.
Marginal spend test
How much further can we push this?
- Treatment
- Spend is increased by a fixed amount — typically 10–20%.
- Control
- Spend stays at business-as-usual levels.
- Read it on
- Δ revenue and Δ profit on the increment only
What to do with it
Increase while you are on the elastic part of the response curve; hold when it flattens.
What it costs you
The effect is smaller than a lift test's, so it needs more volume or a longer window. Results are also specific to the spend range tested and do not extrapolate far.
How a test runs, step by step
This is the shape of the thing you are building towards — two matched groups, one deliberate change, and a gap you can put a number on:
Reading the result
Two matched groups, one spend change. The gap after the change is the incremental effect.
Incremental gross sales
+11.4%
Incremental NGP3
+6.2%
Significance
95%
Illustrative figures. Note the gap between sales and profit — that difference is usually the finding.
Getting there takes six steps.
- Start from a decision, not a channel. “Is Meta prospecting in Germany incremental?” “Should we start TikTok in Sweden?” “Can we reduce branded search in the UK?” A test that is not attached to a decision you would actually make produces a number nobody uses.
- Split the country into balanced groups. Regions are matched on historical sales and other characteristics so the two groups track each other before the test begins.
- Check the split before you spend anything. Run the two groups against each other over historical data with no change applied — an A/A check. If they already diverge, the split is wrong and no result from it will mean anything. This step costs nothing and saves whole tests.
- Choose the campaigns and the direction. Which campaigns are in scope — prospecting, PMax, branded search — and whether spend goes up or down.
- Run it, changing nothing else. The spend change applies to the treatment regions only. Everything else stays on. You are measuring the effect of a specific change, not switching your marketing off.
- Read the outcome. Compare the groups on sales, profit and acquisition cost, check whether the result is significant and stable, and state the options: scale, hold, or cut.
- Act, then feed it back. Update the budget, update the assumptions in your marketing mix model, and write the test down — hypothesis, setup, external events, result, decision. The log is what makes the second year of testing worth more than the first.
Most tests run four to six weeks including a post-treatment observation period. That observation window exists because of carryover: advertising keeps working after it stops, and demand suppressed by a spend cut can partly return later. Cutting the read short overstates the effect of an increase and understates the damage of a decrease. Shorter than four weeks and you are usually measuring noise plus a lag.
What to measure
Most incrementality testing stops at incremental conversions or incremental revenue. That is where it goes wrong, because two channels can produce identical incremental revenue and opposite effects on the business. Three outcomes are worth reading on every test:
Incrementality results
Estimated incremental revenue per channel based on geo-holdout tests run over the last 8 weeks.
By channel
- Gross sales — the incremental revenue the spend change created. The headline, and the least informative of the three on its own.
- Net Gross Profit 3 — the same effect after cost of goods, fulfilment and marketing, with returns deducted first. This is the number that says whether the incremental revenue was worth having.
- Customer acquisition cost — what the incremental new customers cost. It separates channels that grow the customer base from channels that harvest the customers you already had.
The gap between the first and the second is often the whole finding. A channel can deliver a convincing lift in gross sales while selling discounted, return-heavy products at a contribution margin that makes the exercise pointless. If you only read revenue, that channel looks like a winner and you scale it.
This is also why the underlying data model matters more than the experiment engineering. Reading a test on profit requires cost of goods, fulfilment cost by market and route, and a defensible expected return rate per product — before the test starts, not afterwards.
The numbers you will actually report
Three figures come out of a test, and they answer different questions. Worth being precise about them, because they get used interchangeably and they are not interchangeable.
Lift
Lift = (treatment outcome − control outcome) ÷ control outcome
The percentage difference the spend change produced. If treatment regions did €1.34m and matched control regions did €1.20m, lift is 11.7%. It tells you thesize of the effect but nothing about whether it was worth paying for.
Incremental ROAS (iROAS)
iROAS = incremental revenue ÷ incremental spend
The return on the money you actually added, and the number most people mean when they say a test “came back at 2.4”. Note both halves are incremental: the numerator is the measured gap between the groups, not the platform’s attributed revenue, and the denominator is the extra spend, not the whole budget.
This is the figure that tends to shock people. A channel reporting 8.0 in-platform can measure below 2.0 incrementally, because most of what it was claiming was demand that already existed. The comparison, not the absolute number, is what changes decisions:
Reported ROAS vs iROAS
What the platforms claimed, against what a geo test measured.
Illustrative figures. The pattern is the point: the further down the funnel a channel sits, the more of its reported return is demand it did not create. Prospecting barely moves.
Incremental profit — the one that should set the budget
Incremental NGP3 = incremental Net Gross Profit 3 ÷ incremental spend
The same calculation on contribution margin rather than revenue. An iROAS of 2.4 sounds healthy until you know the products behind it carry a 38% return rate and thin margins, at which point the same test reads as break-even. Running the calculation on profit is not a refinement of iROAS; it frequently reverses the decision.
Incrementality rate
Incrementality rate = incremental conversions ÷ attributed conversions
The share of what a channel claimed that it actually caused. Useful as a standing correction factor: if branded search measures at 15%, you can discount its reported conversions accordingly between tests — which is exactly the weighting that causal factor attribution applies, and where its weights come from.
Other ways to test, and when to trust them
Geo testing is not the only option, and it is worth knowing what the alternatives do and do not prove.
- Platform-native lift tests — Meta conversion lift, Google geo experiments and their equivalents. These are properly randomised and genuinely useful, and they have one structural problem: the platform designs the test, measures the outcome and reports the result on its own conversion definition. It is the same party whose performance is being assessed. Use them for direction and for in-platform decisions; treat a platform-run test as weaker evidence than an independent one when the two disagree.
- Audience holdouts and PSA tests — withhold ads from a random slice of the audience, or serve them an unrelated public-service ad instead. Much better statistical power than a geo split because randomisation is at user level, and the right design for CRM, email and SMS. The limitation is that they depend on the platform’s own audience infrastructure, which brings you back to the previous point, and they cannot capture effects that spread beyond the targeted individual.
- Ghost ads — the control group is served an ad they would have seen from a different advertiser, so the comparison is closer to like-for-like than a PSA. Technically elegant, and only available where the platform supports it.
- Before-and-after comparisons — switch something off and see what happens. Cheap, intuitive, and the weakest of all of these, because seasonality, promotions, competitor activity and demand shifts all change at the same time. Without a concurrent control group there is no way to separate them.
Geo testing sits where it does for one reason: the control group is independent of any platform’s infrastructure and any platform’s definition of a conversion, and the outcome is read from your own commercial data. It costs more power than a user-level holdout and buys independence.
In-platform vs independent
Two ways to test. The difference is who runs it and whose data decides the answer.
In-platform tests
Built and run inside a single platform's environment
When to use
Tactical — quick reads within one channel
Independent tests
Platform-agnostic, measured on your own data
When to use
Strategic — cross-channel budget allocation
Quick, directional, inside one channel
CombineVerified, comparable, across channels
Platform tests are properly randomised and useful. They are also designed, measured and reported by the party being assessed — which is why they settle in-channel questions and not budget arguments.
A worked example: four tests, four markets
Aim’n is an activewear brand selling across several international markets, with a large share of budget in Meta. Their question was the one most brands have and few answer: how much of Meta’s reported performance was genuinely incremental, and how much was existing demand or cross-channel overlap being claimed twice.
We ran four incrementality tests across four markets. In each, the treatment regions received no Meta spend while matched control regions continued as usual. The difference between them is the incremental effect — no attribution model involved.
Market 1
2.19×
over-reported
Market 2
2.32×
over-reported
Market 3
2.25×
over-reported
Market 4
3.08×
over-reported
Meta over-reported conversions by an average of 2.46× across all four markets.
The result that matters more than the headline
The obvious conclusion from a 2–3× over-report is “cut Meta”. That conclusion would have been wrong. Measured on epROAS — effective profit ROAS, the return calculated on contribution margin rather than revenue — the campaigns remained above 100% in every market. The spend was genuinely driving incremental profit. Just not remotely at the level the platform claimed.
That is the whole argument for measuring this way. Platform numbers were wrong by a factor of two to three, and the correct response was not to cut but to re-baseline: keep spending, stop believing the reported figure, and set targets against the measured one. A team working from platform ROAS alone would have either over-invested on a false number or, on discovering the gap, cut a channel that was making money.
What it changed
- A true baseline for Meta, per market — including the finding that the over-report was not uniform, ranging from 2.19× to 3.08×
- Scaling decisions made against measured incremental profit rather than reported revenue
- Ground truth to calibrate the marketing mix model with, so the correction persists between tests rather than expiring with the experiment
“Since we started using Dema's incrementality testing, we've gained a much clearer view of what truly drives incremental demand in our paid social. Those insights now guide a significant part of our global spend.”
Kousha Torabi
Co-founder, Ninepine
Building a testing programme
Mature teams do not run tests at random. They tie each one to a decision, sequence them across the year, and keep the operational load manageable. A reasonable progression:
Months 1–6
- One or two lift tests on your largest channels
- A branded search test, if you bid on your own brand at any scale
Months 7–12
- Marginal spend tests on the channels you have now proved are incremental
- New channel tests for platforms you are considering
Year two onwards
- Recurring lift tests to recalibrate, because results decay
- Marginal tests as a standing input to budget planning
- New channel tests whenever a platform or format appears worth evaluating
Standardise so results stay comparable
The value of a testing programme compounds only if this year’s results can be read against last year’s. Fix how you select treatment and control regions, fix a minimum duration, and fix the metrics. Then document, for every test: the hypothesis, the market, the campaigns and spend change, the timeline, anything unusual happening in the market at the time, the result, and the decision you took.
That record becomes the most valuable measurement asset you own — a body of causal evidence about your own business that no vendor benchmark can substitute for.
How test results calibrate your MMM
A finished test produces a ground-truth point: “Meta in Germany produced this much incremental gross sales at this spend level”, or “branded search in the UK is not incremental”. Those points are what a marketing mix model should be anchored to.
Given them, the model can:
- Constrain channel coefficients to ranges the evidence supports
- Avoid over-crediting channels with strong correlation and weak causation — the failure mode MMM shares with attribution
- Be defended in a budget meeting, because part of it has been measured rather than fitted
The same results have a second use. They are the evidence base for deciding how much of each attribution source to believe — causal factor attribution applies explicit weights to ad platform and multi-touch claims, and those weights should come from your own experiments rather than from an assumption. Run more tests and both the model and the attribution get sharper. Unified measurement is the practice of running all of these together and reconciling them.
A calibrated model gives you response curves: how return changes as spend rises on each channel. That is the output a budget is actually set from, and it is only trustworthy at the points where an experiment has pinned it down. Marginal spend tests are what anchor the steep part of the curve; without them, the shape beyond your historical spend range is extrapolation.
What incrementality testing cannot do
Incrementality is the strongest evidence available in marketing measurement, and it is still evidence rather than truth. Anyone selling it as a single always-correct number is overselling it.
- It costs time and money. A test means running a suboptimal spend plan in part of a market for several weeks. That is a real cost, and it is the price of the evidence.
- You cannot test everything. Each test occupies a market for weeks, so tests are a scarce resource to be spent on the decisions that matter most. This is precisely why MMM is needed alongside them.
- Small markets may not work. If a market has too little volume, no test design will separate a real effect from noise. Better to know that in advance than to run a test and interpret a null result as a finding.
- Spillover blurs the boundary. A geo split assumes the two groups are independent, and they are not entirely. Customers move between regions, word of mouth crosses borders, and national media and platform optimisation do not respect your split. Spillover generally makes an effect look smaller than it is, so a positive result stays trustworthy — but a null result deserves a second look.
- External shocks contaminate results. A competitor’s campaign, a heatwave, a stockout, a PR event. Some of this can be controlled for, and some can only be noted and taken into account when reading the result.
- Results are specific to their setting. A finding holds for that market, that period, that spend level and that creative. It is not a permanent fact about the channel, which is why recurring tests exist.
- Some things are out of scope. Geo-based testing suits large paid platforms. CRM, email and SMS need audience-level holdouts instead, and offline media on its own is difficult to isolate this way. It is also worth being clear that this tests a specific spend change on specific campaigns — it is not switching your entire marketing engine off to see what happens.
The honest framing is the useful one: incrementality gives you better causal evidence for decisions than anything else available. It does not give you a number that is always right.
What different teams get from it
Incrementality is usually bought by marketing and valued most by everyone else.
Founders and C-suite
Neutral evidence when platform numbers and internal opinions collide. Strategic questions — is this channel worth its budget, is this new platform adding anything — get answered rather than argued.
CMOs and heads of e-commerce
“Let's test it” replaces “let's argue about it”. You learn where you can safely cut and where to double down, and build a measurement practice that outlasts any single platform's reporting changes.
Performance and growth marketers
Proof for the things you already suspect but cannot show from a dashboard, and a defensible basis for asking for more budget — which is a much stronger position than a high ROAS figure.
Finance and CFOs
The conversation moves from ROAS to incremental profit and acquisition cost, and it becomes visible where marketing adds profitable volume versus where it moves sales that were coming anyway.
Frequently asked questions
Incrementality testing measures the causal effect of marketing by comparing regions where you change spend against regions where you leave it unchanged, then reading the difference in business outcomes. Because the only thing that differs between the two groups is the spend, the gap between them is what the spend caused. It is the only measurement method that observes what happens when marketing stops, which is why it is treated as the ground truth that other methods are calibrated against.
Attribution divides credit among conversions that already happened, so it can only describe what it observed and inherits every gap in that observation. Incrementality testing changes spend in some regions and not others and measures the difference in actual outcomes, which answers whether the spend caused anything at all. Attribution can tell you that a customer touched three channels; only an experiment can tell you whether they would have bought anyway.
Most tests run four to six weeks, including a post-treatment period to catch delayed effects. The right length depends on your purchase cycle and volume: a high-frequency, high-volume category can read a result faster than a considered purchase with a long journey. Running shorter tests to get answers sooner is usually a false economy, because you end up measuring noise plus a lag and then treating it as a finding.
Enough volume in the market to detect an effect, the ability to change spend by region for the campaigns in scope, and — if you want to read the result on profit rather than revenue — cost of goods, fulfilment costs and an expected return rate per product. That last requirement is the one teams discover late. Reading a test on contribution margin is only possible if the cost model already exists.
Sometimes, and it varies enough by market that it has to be tested rather than assumed. Brand bidding intercepts people who are already looking for you, so a share of those conversions would have arrived through organic results at no cost. But the share is not 100%: competitors and resellers bidding on your name can take real demand if you concede the auction. A branded search test measures the actual trade-off in your markets, which is usually the single highest-value test a brand can run first.
iROAS is incremental return on ad spend: the incremental revenue a test measured divided by the incremental spend that produced it. Ordinary ROAS divides all attributed revenue by all spend, so it includes conversions that would have happened without the advertising. The gap between the two is often large and it is not uniform — branded search and retargeting typically collapse the most, because they intercept demand that already existed, while prospecting changes least. That is why a channel can report 8.0 in-platform and measure under 2.0 in a geo test. The version worth budgeting on goes one step further and runs the same calculation on contribution margin rather than revenue.
It varies by channel and by market, which is itself the point — there is no single correction factor to apply. In four tests we ran with the activewear brand Aim'n, Meta over-reported conversions by 2.19×, 2.32×, 2.25× and 3.08× in four different markets, averaging 2.46×. The spread matters as much as the average: a brand applying one blended discount across markets would have been wrong in three of the four. And the honest footnote is that Meta was still profitable in every market once measured on profit rather than revenue, so the correct response was to re-baseline the targets, not to cut the channel.
They are properly randomised and genuinely useful, with one structural caveat: the platform designs the test, measures the outcome and reports it against its own conversion definition, while being the party whose performance is under assessment. That is not an accusation of bad faith — it is a conflict of interest worth pricing in. Use them for in-platform decisions and direction. When a platform-run test and an independent geo test disagree, weight the independent one, because its control group and its outcome data do not belong to the party being measured.
They work much better together. Incrementality gives precise causal answers to specific questions, but each test occupies a market for weeks, so you cannot cover every channel continuously. MMM gives continuous coverage of the whole mix, but it is a model and can drift. Tests anchor the model to measured facts, and the model covers the ground between tests. Neither substitutes for the other.
That the test could not distinguish the effect from random variation — not that the channel does nothing. The distinction matters commercially, because treating an underpowered test as proof of no effect is how teams cut channels that were working. If the spend contrast was small, the market thin or the window short, the correct conclusion is that the test was not able to answer the question, and the design needs changing before you conclude anything.
Not well with geo-based designs, which is the approach best suited to large paid platforms. CRM channels like email and SMS need holdout designs at the audience level rather than the geographic level, and offline media on its own is hard to isolate geographically because its footprint rarely matches a clean regional split. Being clear about that boundary is part of running a credible programme.
How Dema runs incrementality tests
Dema automates the parts that make testing hard to sustain: geo selection and matching, synchronising the spend change with the ad platforms, and reading the result on gross sales, Net Gross Profit 3 and new-customer acquisition cost from your own commercial data. Results feed the marketing mix model and the attribution weights automatically, so each test makes the ongoing measurement sharper rather than sitting in a slide deck.