Growth that compounds

Marketing mix modeling is a calibration problem

Marketing mix modeling can organise an aggregate budget, but it cannot certify causality by itself. Calibrate the model before trusting the allocation.

1,163 words 5 min read 4 references  readers

Management summary

Marketing mix modeling is often asked to do two jobs at once: describe how aggregate demand moved and prove which channel caused the movement. Those are different jobs. A model can be useful for planning while remaining vulnerable to seasonality, targeting, spend response and the choice of transformations. The practical question is therefore not whether to use marketing mix modeling or an experiment as if they were rival religions. It is what each instrument can identify, how the model is calibrated, and which decision is allowed to rely on the output. This article compares attribution, aggregate modeling and holdout tests, then turns the comparison into a calibration log that a commercial team can maintain without pretending the model has seen a counterfactual.

Keywords: Marketing mix modeling · Incrementality testing · Advertising measurement · Budget allocation

On this page

Marketing mix modeling is a calibration problem, not a dashboard problem. It can organise an aggregate budget decision, show how demand and spend moved together, and expose a set of scenarios. It cannot certify causal lift just because the regression is sophisticated.

The distinction is easy to lose because every instrument returns a number. A multi-touch attribution model assigns credit to observed touches. A marketing mix model relates aggregate outcomes to aggregate inputs. An incrementality test asks what changed relative to a counterfactual. The numbers look comparable only after the question has been silently changed.

Which three distinct measurement questions do attribution, experimentation, and MMM answer?

The first discipline is to name the output before choosing the model. If the decision is whether a channel created demand that would not otherwise have existed, the model must contain a credible counterfactual. If the decision is how to plan a portfolio under several demand scenarios, an aggregate model may be the right instrument even when it cannot identify causal lift on its own.

InstrumentWhat it observesStrongest defensible outputInvalid leap
Touchpoint attributionRecorded exposures and conversionsWhich touches are present in the observed pathThose touches caused the conversion
Marketing mix modelingAggregate outcomes, spend and controls over timeA scenario relationship under stated model assumptionsThe coefficient is experimental lift
Holdout or lift testTreated and untreated units under a designed interventionIncremental outcome under the tested conditionsThe result transfers unchanged to every channel
Calibrated portfolio viewModel scenarios plus experiments and operating boundsA decision with a stated confidence and downsideOne dashboard is the truth

Table 1What each measurement instrument can carry

The instrument is not judged by whether it produces a number. It is judged by the question that number can answer.

Source: Author's synthesis of the cited measurement literature.

View exhibit page

This is the narrower claim behind the incrementality illusion. The problem is not that an observed association is useless. It is that an organisation starts making a counterfactual decision from an instrument that never observed the counterfactual.

Why do marketing mix models require continuous experimental calibration against lift tests?

An aggregate model sees time. Time carries seasonality, promotions, product launches, macroeconomic shocks and the company’s own decision to spend more when demand looks promising. The model must separate those movements with data and assumptions that are never perfectly observed.

The transformations matter too. Adstock, saturation, lag and interaction terms are not decoration. They are claims about how exposure becomes demand. Different assumptions can produce different allocations while fitting the past in a similar way.

Calibration is the discipline that keeps this from becoming a debate about whose chart looks more plausible. It uses experiments, external shocks, stable control series, known pricing changes or other observations to constrain the model. It does not make the model experimental. It tells the reader which parts of the model have an anchor outside the model.

Gordon and colleagues show why this matters when different advertising measurement approaches are compared against field experiments. Blake, Nosko and Tadelis show how paid search can capture inframarginal buyers. Lewis and Rao show how noisy outcomes make return measurement expensive. The lesson is not to discard every model. It is to stop treating model fit as proof of lift.

Model elementAssumptionExternal anchorIf the anchor moves
Baseline demandThe non-media demand path has this shapeControl series, category data or known interruptionRe-estimate the base before reallocating
CarryoverExposure persists for this longDelayed response in an experiment or prior with a stated basisWiden the scenario range
SaturationAdditional spend produces less response after this pointSpend variation with an independent shockDo not use the point estimate as a ceiling
Channel interactionTwo channels reinforce or substituteDesigned test or a documented mechanismKeep the interaction as a scenario, not a fact
Incremental liftThe allocation reflects causal contributionHoldout, ghost-ad or lift resultMark the model coefficient as uncalibrated

Table 2The calibration log

A model becomes more credible when its assumptions and outside anchors are visible in the same row.

Source: Author's worksheet, informed by the cited advertising measurement studies.

View exhibit page

How can commercial leadership apply MMM coefficients without econometric overreach?

Start with the portfolio question. Which channels need a directional bound? Which channels have enough scale for a designed holdout? Which decisions are defensive and should be governed by a budget ceiling rather than a claim of incremental demand?

Then show the range. A model that produces one allocation without a sensitivity view has hidden its uncertainty. Vary the baseline, the carryover and the saturation assumptions. If the recommended portfolio changes when a reasonable assumption moves, the decision is sensitive to the assumption. That is an output worth knowing.

The next step is to connect the model to a test. A holdout can calibrate a channel or a family of channels. It may not transfer to every market or period, but it gives the aggregate model an external point. The evidence over anecdote worksheet is useful here: write down what would count as evidence before the allocation is shown.

Finally, separate the model’s scenario from the operating decision. The model can say that a portfolio looks better under one set of assumptions. Leadership still has to decide how much uncertainty it is willing to buy, what cash payback is acceptable and which learning matters more than this quarter’s apparent efficiency.

Why is executive econometric governance distinct from dashboard monitoring?

The dashboard is a view. Governance is the rule that says what happens when the view conflicts with a test, a margin threshold or a customer constraint. The funnel bottleneck nobody’s measuring shows the same pattern outside advertising: the visible metric is often the output, while the unmeasured interval decides whether the output can be trusted.

A marketing mix model is worth keeping when it makes assumptions explicit, supports scenarios, accepts outside calibration and changes a decision. It is not worth keeping as a decorative source of precision. The right question at the next review is not “What did the model say?” It is “Which part of the model has been tested, which part is a scenario, and what decision is allowed to rely on each?”

References

  1. Gordon, B. R., Zettelmeyer, F., Bhargava, N., & Chapsky, D. (2019). A comparison of approaches to advertising measurement: Evidence from big field experiments at Facebook. Marketing Science, 38(2), 193–225. https://doi.org/10.1287/mksc.2018.1135
  2. Blake, T., Nosko, C., & Tadelis, S. (2015). Consumer heterogeneity and paid search effectiveness: A large-scale field experiment. Econometrica, 83(1), 155–174. https://doi.org/10.3982/ECTA12423
  3. Lewis, R. A., & Rao, J. M. (2015). The unfavorable economics of measuring the returns to advertising. The Quarterly Journal of Economics, 130(4), 1941–1973. https://doi.org/10.1093/qje/qjv023
  4. Johnson, G. A., Lewis, R. A., & Nubbemeyer, E. I. (2017). Ghost ads: Improving the economics of measuring online ad effectiveness. Journal of Marketing Research, 54(6), 867–885. https://doi.org/10.1509/jmr.15.0297

Pass it on

Share this essay

If it was useful to you, it is probably useful to someone on your team.

Download as PDF

A complete document: title page, contents, sources, and the citation on the last page.

Sinan Isoglu

About the author

Sinan Isoglu, MBA (Quantic)

Commercial growth leader, lecturer and doctoral researcher

Sinan Isoglu is a commercial growth leader, lecturer and doctoral researcher. His work spans go-to-market, pricing and revenue operations; his doctoral research at EM Normandie examines sales and marketing integration after cross-border M&A. He lectures on marketing and growth at IU International University of Applied Sciences.

Credentials

  • Doctoral researcher, EM Normandie Business School
  • MBA, Quantic School of Business and Technology
  • Lecturer, IU International University of Applied Sciences

Writes on

  • Go-to-market
  • Pricing
  • Revenue operations
  • AI in commerce
  • Cross-border growth

The track

The work behind this question.

This piece sits in the commercial track: the operating problems behind growth, pricing and revenue systems.

Comments

Join the thinking.

Comment on the piece, or select a passage above to quote it directly.

Leave a comment

Comments are read and approved personally before they appear. Your name and comment are stored for publication. See the Privacy note.