Clean, weak and non-identification

A converged model is not an identified one. This page defines the three identification states an MMM coefficient can be in, explains why the middle state is the dangerous one, and argues that when the data do not support the question, the defensible move is to ask a weaker question rather than deliver a confident answer to the original one.

The three states

Identification is a property of the likelihood: how sharply the data distinguish one parameter value from its neighbours in the direction that matters.

StateLikelihood in that directionPosterior behaviourSampler behaviour
Cleanly identifiedCurved; the data prefer a regionConcentrates as data accumulate; stable under reasonable prior changesClean
Weakly identifiedNearly flat; many values fit almost equally wellFinite, plausible-looking, dominated by the prior and the functional formUsually clean
Non-identifiedExactly flat; the data cannot distinguish values at allReturns the prior, possibly reshaped by the transformOften still clean

The third column is the trap. All three states can produce a posterior with good R-hat, adequate effective sample size and no divergences, because MCMC diagnostics assess whether the sampler explored the posterior, not whether the data built it. A Bayesian model composes prior and likelihood into a proper posterior even when the likelihood contributes nothing in some direction. That is a feature for honest uncertainty propagation and a hazard for anyone reading the interval as evidence.

The boundaries are not symmetric in kind. Exact non-identification is a mathematical property, and AMMM3’s FE and CRE estimability screens catch its arithmetic form, zero within-unit variation, before fitting. The boundary between weak and adequate is not a property the data announce. It is a judgement about the decision at stake: variation that supports “this channel contributes something” may not support “move budget from search to this channel”. See Always-on and low-variance channels for the full treatment of that boundary and the screens that locate it.

Why weak identification is the dangerous case

Non-identification announces itself if you look: the posterior equals the prior, and the estimability screens refuse the worst cases outright. Weak identification produces output that passes every routine inspection.

The posterior is finite, unimodal and plausibly located. With a LogisticSaturation and a positive-support beta, a weakly identified channel still returns a non-zero contribution, a finite ROAS and an interval. The number is not random. It is the prior propagated through the transform, and it looks exactly like a result.

The interval understates the real uncertainty. The reported interval is conditional on the model: the transforms, the control set, the priors. In the weak state those conditioning choices carry most of the weight, so the interval measures the prior’s confidence, not the evidence’s. The honest uncertainty about the channel’s effect is far wider than the printed interval, and nothing on the output page says so.

The point estimate can be confidently wrong, and the error is systematic. In a weak direction the fit is settled by whatever correlated structure is available: the prior’s location, the saturation shape, a co-moving channel, a trend term. These push the estimate in one direction consistently across refits. The result behaves like a bias, not like noise, so refreshing the model or adding data from the same schedule reproduces the same wrong number with the same false precision. The baseline and media trade-off and the price-inversion example in Price and promotion in an MMM show the same mechanism: predictively excellent fits carrying signs and magnitudes that are artefacts of the specification.

“Directional” is not a weaker version of correct

The common rationalisation is that a weakly identified estimate is still “directional”: imprecise but pointing the right way, and better than telling the client nothing. The rationalisation fails on its own terms.

In the weak state the direction is exactly what the data did not determine. The sign came from the prior’s support, the transform’s shape or a correlated regressor, so “directional” promotes the least data-driven feature of the output to the headline. A positive-support channel prior cannot produce a negative coefficient no matter what the data say; reading its positive posterior as directional evidence is circular.

The client cannot see the difference. A wide interval honestly reported is detectable by the reader, who can discount it. A confident number from a weak fit is indistinguishable on the page from a confident number from a clean fit. Delivering it converts a visible absence of evidence into an invisible assumption, and the error surfaces later as a budget decision that does not replicate.

The defensible alternative: weaken the question

When identification is weak, the data usually do support something. The professional move is to find the strongest claim the evidence carries and report that, rather than the claim the client asked for.

  • Narrow the estimand. Group the weak channel with related channels and report the group’s contribution, stating that the group result does not licence reallocation within the group.
  • Downgrade the claim class. Report a conditional model allocation with its prior sensitivity shown, instead of an incremental effect. The agency workflow claim classes exist for exactly this distinction.
  • Answer the decision, not the coefficient. If the decision is “keep or cut”, a bounded statement (“under all priors we tested, the contribution stays below X”) can be supportable when a point estimate is not.
  • Propose the design that would answer the original question: created spend variation, a geo holdout or an incrementality test, which can then enter supported estimators as calibration.

“The data do not support the channel-level question yet; here is what they do support, and here is what would answer it” is a deliverable. It is also the only one of these outputs that remains true after the next refresh.

How to tell which state you are in

  1. Prior sensitivity is the direct test. Refit under a defensibly different prior; a posterior that moves materially with the prior is weakly identified in that direction. The pipeline’s prior sensitivity surface and parameter dependence reports operationalise this.
  2. The FE and CRE estimability reports flag zero within-unit variation as an error and low variation share, high VIF and high condition number as warnings. A warning does not block the fit; it obliges the sensitivity check above. Pooled fits have no such screen, so absence of a warning there is absence of a check, not evidence of identification.
  3. Pre-fit data screens predict most weak cases before sampling: see How much data does an MMM need?.
  4. Predictive performance does not discriminate. Weakly identified models often predict well, because prediction rewards the fitted sum, not the allocation. Do not use holdout error to overrule a failed sensitivity check; see Model comparison for econometricians.