How much data does an MMM need?
The honest answer is that the binding constraint is identifying variation, not row count. A dataset with four years of weekly observations can contain less information about channel effects than a dataset with two years, if the four years were spent moving every budget together. This page gives the practical arithmetic and the screens to run before promising a client anything.
Rows are not information
The model estimates each channel’s amplitude, carryover and curvature from the contrast between periods of high and low transformed exposure, after the baseline, seasonality and the other channels have claimed their share of the same movements. What counts is therefore:
- how much each channel’s transformed spend varies within the sample
- how much of that variation is shared with other channels, trend and seasonality
- how many independent movements the sample contains, not how many rows
Two hundred weeks of spending that tracks a shared campaign calendar contain few independent movements. Fifty-two weeks with staggered flighting across channels can contain more. Aggregation compounds this: national weekly totals average away the regional and daily contrasts that would otherwise identify the effects.
What history length buys
Length matters most for the nonlinear transform parameters, because each is estimated from a specific kind of contrast.
| Parameter | Identified by | Consequence of a short sample |
|---|---|---|
| Adstock decay | Spend changes followed by observable outcome decay, within l_max | Few on-off cycles means the decay posterior stays close to its prior |
| Saturation curvature | Spend observed at materially different levels | Spend at one operating point leaves the curve prior-shaped away from it |
| Seasonality | Repeated cycles | One observed year cannot separate annual seasonality from trend or one-off events |
| Channel amplitude | Everything above, jointly | Inherits the weakest of the three |
Two annual cycles is a practical floor for separating annual seasonality from trend, and three is materially better. For carryover, the number of distinct spend pulses matters more than calendar time: a channel that changed level twenty times in eighteen months is better identified than one that changed four times in four years.
Weekly data is the usual resolution. Daily data multiplies rows but adds
day-of-week structure and noise, and adstock horizons measured in days make
l_max large; aggregate deliberately rather than assuming finer is better.
Monthly data usually cannot support adstock estimation at all, because most
carryover completes within the observation interval.
What a geo panel buys and costs
A balanced geo panel multiplies the identifying variation when spend timing or level differs across units, and the FE and CRE presets are built to use exactly that within-unit variation. Ten geos with independently timed flights can identify effects that a national aggregate of the same activity cannot.
The costs are contractual and statistical. The panel must be balanced on unit and date, with every channel varying within at least one unit; see the panel data layout and the FE and CRE pages. National media that is identical in every geo gains nothing from the panel, because it has no within-unit variation to use. And FE removes only time-invariant unit differences; it does not create identification for time-varying confounding or common shocks.
Screens to run before committing
Run these on the candidate dataset before the engagement promises channel-level results.
- Per channel, the coefficient of variation of spend, the number of distinct spend levels and the longest constant run. A channel that never moves cannot be estimated at any sample size; see Always-on and low-variance channels.
- Pairwise correlations of transformed spend between channels and with a linear trend. Shared calendars show up here, and the model will allocate shared variation by prior and functional form.
- The number of observed annual cycles for any seasonality claim.
- For a panel, whether spend timing actually differs across units, and whether the panel is balanced.
- For FE and CRE, the estimability report’s within-variation shares, VIF and condition number after fitting starts; the screens above predict most of what it will say.
If the screens are poor, the defensible responses are to narrow the estimand (group channels, report aggregates), to create variation prospectively (flighting, geo splits), or to say that the data do not support channel-level answers yet. More rows of the same schedule will not change the answer.
The finite-sample connection
Short samples are also where the Bayesian machinery earns its keep and where its cost appears. The posterior is exact at any sample size, conditional on the model and prior, but in a short or low-variation sample the prior carries more weight, so the prior-sensitivity comparisons in the agency workflow are the direct check on whether the data or the prior produced the estimate. There is no sample size at which that check becomes unnecessary, and no sample size below which modelling is impossible; the question is always which estimands the available variation supports.