MMM vs incrementality testing
An incrementality test is a controlled experiment: suppress or boost spend in some markets, hold others as control, measure the causal difference. An MMM is a statistical model over history. The experiment is ground truth for one question at a time; the model covers every channel, every week. Mature stacks use both.
By Oliver Wakefield-Smith · Updated 2 August 2026 · Every number on this page resolves to a primary source in the ledger below.
What a geo-lift test actually does
Pick a set of test markets, change spend there (dark it, or surge it), keep matched markets as control, and compare outcomes against what the control markets imply the test markets would have done. Platforms like Haus productize the market matching, the power analysis, and the readout.1 The output is a causal lift estimate with a confidence statement: the closest marketing gets to a clinical trial.
Why they need each other
The experiment's power is credibility; its cost is scope. One test answers one question over one period, at the price of deliberately unbalanced spend for weeks. You cannot run fifty simultaneous holdouts, and you cannot re-run last year. The model's power is coverage, every channel decomposed continuously; its weakness is that it is inference from observational data, only as good as the variation it saw.
Calibration closes the loop: experiment results enter the model as informative priors or constraints, anchoring the statistical estimates to measured truth. Meridian was built with this workflow in mind, and PyMC-Marketing documents lift-test calibration as a first-class feature.23 A model calibrated to two or three experiments per year earns a different level of trust than one that has never been checked against reality.
Which questions belong to which instrument
Only a test can answer: does this channel work at all for us; how much of branded search is cannibalized; did that big campaign cause anything. Only a model can answer: how should next quarter's budget split across nine channels; where is each channel on its saturation curve; what did the whole mix contribute last year. If a vendor claims one instrument answers both lists, they are selling the instrument, not the answer.
Cadence: the working pattern among sophisticated teams is one to two calibration experiments per quarter, rotated across the largest or most contested channels. Big channels first; the experiment's cost is proportional to spend held back, but so is the cost of being wrong about the channel.
Continue to triangulation, where the two instruments plus platform data become one operating system.
Source ledger
Every numeric claim on this page resolves to one of these primary sources.
- [1]Haus public sitehttps://www.haus.io/retrieved 2026-08-02
- [2]Google Meridian documentationhttps://developers.google.com/meridianretrieved 2026-08-02
- [3]PyMC-Marketing documentationhttps://www.pymc-marketing.io/retrieved 2026-08-02