Methodological benchmark protocol v1
Status and authority
| Field | Frozen value |
|---|---|
| Protocol identifier | AMMM3-METHOD-1 |
| Version | 1.0 |
| Effective date | 11 September 2026 |
| Status | Frozen methodological baseline |
| Scope | Technical capability, evidentiary discipline and reproducible agency workflow |
| Governing guide | A Bayesian workflow for agency teams |
| Comparative results | None established by this document |
This protocol assesses whether a tool, configuration and operating procedure can support a stated agency use with inspectable evidence. It does not rank products by predictive scores, runtimes, numerical recovery results or a weighted feature total. Numerical diagnostics may be retained as evidence for an individual acceptance decision; they are not a comparative leaderboard.
The freeze fixes requirements, claim boundaries, case families and adjudication rules. It does not certify AMMM3 or any competitor. Each implementation review must also freeze its own case specification and software versions before examining outputs. No executable cross-product case pack is supplied by this version.
Material changes to these rules require a new versioned protocol page with a reason and a description of affected decisions. Preserve this version as the basis of existing reviews. Corrections that change no requirement may be made only with a dated amendment entry. A current implementation failing a rule is not grounds for silently changing that rule.
Unit of assessment
Assess a named product version, estimator, configuration, public entry point, case and intended use. Separate the statistical engine from orchestration, reporting and human review. A capability may be native, supplied by a named extension, or performed through a documented manual procedure. Record which mechanism supplies it and retain the evidence for the complete workflow.
The comparison set comprises AMMM3, PyMC-Marketing, Meridian, Robyn and named lightweight commercial workflows. A commercial entry must identify its product or actual spreadsheet/notebook procedure, version, accessible inputs and exportable evidence. Do not assess an unspecified category as if it were one product. Google’s LightweightMMM would be a separate named entry if included.
Each product may use its documented operating procedure. Common cases must share the decision question, input snapshot, outcome units, relevant history, calibration evidence and evaluation split. Match model equations, transforms and priors where that is feasible and material. Record differences in lag normalisation, likelihood, regularisation, constraints, scaling or pooling. Do not label different statistical models as equivalent because their APIs accept the same dataset.
A tool may validly decline a case outside its supported contract. That limits its scope; it is preferable to silently accepting an unsupported use. Assess AMMM3 against the same rule, including its named RE gate and FE/CRE limits.
Freeze each case before inspection
The case owner must retain the following specification before reviewing results.
- Protocol version, case identifier, intended claim and decision quantity.
- Data or simulation source, content hashes, schema, units and transformations.
- Known design facts, unavailable information and the reason the case challenges the intended use.
- Model versions, equations or their documented representation, priors or penalties, calibration mechanism and public execution steps.
- Evaluation unit, split, lag-history treatment and information available at prediction time.
- Diagnostic profile, practical decision tolerance and expected handling of missing or failed checks.
- Required alternatives, expected acceptance behaviour and a test of relevant failure handling.
- Required retained artefacts, replay tolerance, reviewer and amendment history.
Use a reviewer-held challenge case or a held-out segment where practical. Exploratory work can inform a later case specification, but must be labelled as exploratory. If results motivate a new tolerance, specification or case, version the change and retain the earlier outcome. Do not overwrite inconvenient cases or present a tuned exercise as untouched validation.
Evidence states
Use these states for each requirement. They describe evidence for the exact assessment scope, rather than overall product quality.
| State | Meaning |
|---|---|
| Demonstrated | Retained outputs and an inspectable procedure support the required behaviour, including relevant failure handling |
| Documented only | A primary source describes the capability, but the review has not demonstrated it |
| Failed | An executed case contradicts the acceptance rule |
| Not assessed | Access, evidence or an executed check is missing |
| Outside supported scope | The named implementation explicitly excludes the use and identifies that boundary |
Also record the delivery mechanism as native, extension or manual. Documentation alone cannot promote a requirement to demonstrated. Manual review can satisfy a procedural requirement with an inspectable signed record; it cannot prove that a computation, API guard or automatic block exists.
For a proposed use, all applicable mandatory requirements must be demonstrated. A mandatory requirement marked failed, documented only, not assessed or outside supported scope blocks acceptance for that use. A restricted use must meet every requirement applicable to the narrower claim and explicitly exclude the original unsupported claim. There is no compensating score that lets a strong report cancel a failed evidence check.
Mandatory requirements
The following requirements are normative. Their identifiers must appear in the review record with scope, evidence state and rationale.
M1 Decision and identification contract
Required evidence. A defined outcome, population, baseline, intervention or scenario contrast, horizon and decision quantity. Retain the causal diagram or assumption account, sources of relevant variation, and the classification of each substantive claim from the governing guide.
Acceptance rule. The claimed interpretation follows from the stated design and assumptions. A causal use explains confounding, support, interference, measurement and transport threats. A conditional use discloses those limits.
Failure case. Supply strong predictive fit with unresolved confounding. The workflow must withhold causal approval. No method is required to discover an unobserved confounder from the observed data alone; it must avoid certifying assumptions that those data cannot verify.
M2 Data and likelihood contract
Required evidence. Fixed input and transformation records, row and unit alignment, missingness treatment, scaling, likelihood and observation dependence. Retain the model’s lag and baseline construction, including preprocessing fitted outside the main estimator.
Acceptance rule. The executed model matches the declared specification. Prediction and calibration use compatible units and indexing. Held-out outcomes do not influence training transformations or tuning presented as independent validation.
Failure case. Introduce a unit mismatch, missing value treated as zero, misaligned channel, or future outcome leakage. The workflow must reject the input or expose the problem and block the affected claim. A justified repair requires a new recorded input state.
M3 Priors, penalties and calibration
Required evidence. Resolved priors or regularisation, scale interpretation, prior predictive checks where applicable, external evidence provenance and the mapping from experimental contrast to model response. Identify whether external information enters as a prior, likelihood or penalty.
Acceptance rule. Constraints represent stated assumptions. Calibration acts on the intended quantity, includes its relevant uncertainty and addresses transport and overlapping evidence. Inspecting the graph contribution or a controlled perturbation can demonstrate coupling; narrower intervals alone cannot. Numerical prior changes require parameter identity, scale and a reason.
Failure case. Use conflicting, mis-scaled, overlapping or inapplicable calibration evidence. The workflow must expose the conflict and withhold an unsupported calibrated claim. It must not count reused calibration evidence as independent validation.
Calibration is mandatory when a workflow claims calibrated inference. A model without external calibration may be assessed for other uses, with its evidence limits stated. Absence of an experiment does not automatically forbid an observational causal argument, but M1 still applies in full.
M4 Computation and uncertainty
Required evidence. Method-appropriate diagnostics, warnings, stopping rules, and uncertainty for the estimands used. For MCMC, retain chain behaviour, effective sample sizes, Monte Carlo precision and sampler pathologies. For regularised or searched models, retain the search/selection procedure and the basis of any resampling or other uncertainty estimate.
Acceptance rule. Required diagnostics are usable and material failures are resolved before affected estimates are interpreted. Identify uncertainty that is included and omitted, especially conditioning on selected hyperparameters, model choice, calibration and measurement. A posterior, bootstrap distribution and range across candidate models are not interchangeable uncertainty objects.
Failure case. Retain an incomplete fit, missing diagnostics, failed sampler check or degenerate resampling result. The review must not infer reliability from successful execution, a narrow interval or a polished report.
M5 Adequacy and predictive use
Required evidence. Relevant outcome and residual checks, an evaluation aligned to the claimed prediction task, and a simple reference prediction. Document time or geographic dependence, known future inputs, carryover history and any model comparison measure’s likelihood basis.
Acceptance rule. Material mismatch is investigated and unresolved limitations constrain use. Predictive claims require relevant held-out evidence. Comparison uses compatible targets and acknowledges uncertainty. Information criteria are optional tools, not compulsory proof of methodological quality.
Failure case. Present good in-sample fit with a held-out mismatch, an inadequate holdout, or incompatible likelihood bases. The workflow must expose the limit. It must not claim causal validation from predictive accuracy, or future forecasting validity from ordinary observation-level LOO alone.
For a use that makes no predictive claim, retain relevant model-adequacy checks and explicitly exclude predictive validation if it cannot be assessed.
M6 Identification and model stability
Required evidence. An estimand-specific assessment of useful variation, parameter dependence and completed sensitivity checks against the most material plausible alternatives. Retain alternative fits, unsuccessful attempts and the reason each alternative was considered.
Acceptance rule. Statistical uncertainty is separated from Monte Carlo error and structural non-identifiability. The reviewed conclusion meets the predeclared practical tolerance across credible alternatives, or the permitted use is narrowed. Likelihood invariance or an applicable exact dependency is needed for a structural non-identifiability claim.
Failure case. Present a correlated channel split that changes under plausible priors or baseline choices while total prediction remains similar. The workflow must expose the instability. Prior concentration, a sensitivity plan or a passing raw-design screen cannot substitute for completed evidence.
M7 Planning and optimisation
Required evidence. A baseline plan, budget units, scenario horizon, carryover treatment, coordinate mapping, constraints, objective and uncertainty for the proposed plan. Identify the observed and calibrated support for the spend changes and the assumptions about non-media drivers.
Acceptance rule. The evaluated plan matches the requested plan, respects constraints and retains its assumptions. State whether optimisation uses a point estimate, a posterior expectation or another objective. Review the plan under relevant uncertainty and model alternatives. A synthetic future path must not be described as an identified intervention by construction.
Failure case. Change budget units or horizon, omit lag context, request an unsupported unit panel or extrapolate beyond support. The workflow must detect a contract violation or explicitly restrict the claim. A deterministic optimum must not be presented as a complete account of decision uncertainty.
This requirement applies when scenario planning or optimisation is claimed. Unsupported optimisation does not invalidate a narrower modelling use.
M8 Reproducibility and evidence integrity
Required evidence. Input identity, resolved model and environment, calibration provenance, fitted state, diagnostics, alternatives, scenario specification, replay steps and named review decision. Retain the link between each report claim and its supporting artefacts.
Acceptance rule. Another analyst can replay a material result within the predeclared tolerance. Distinguish replay from a fresh stochastic fit and from replication on independent data. Changed data, configuration or fitted state must not silently inherit an earlier approval. Hashes need a retained reference and do not establish statistical validity or tamper-proof governance alone.
Failure case. Remove a required artefact, substitute a fitted state or alter a scenario after review. The workflow must expose the missing or changed basis and require a new review before the affected output is used externally.
A manual check can meet this operating requirement if its procedure and result are retained. It cannot support a claim of automatic integrity enforcement.
Case families
Every assessment must specify concrete inputs and expected handling for the applicable families below. These are required case designs, not datasets or results already supplied with AMMM3.
| Family | Purpose | Required observation |
|---|---|---|
| Ordinary aggregate series | Establish a usable end-to-end path | Explicit model, reviewed uncertainty and replayable result |
| Correlated media | Challenge channel separation | Instability is exposed and aggregate claims are assessed separately |
| Confounded media timing | Challenge causal interpretation | Predictive fit does not receive unsupported causal approval |
| Noisy or revised measurement | Challenge data handling | Repairs and measurement assumptions remain visible |
| Relevant and conflicting experiments | Challenge calibration and transport | Contrast mapping and conflicts are reviewed without double counting |
| Geographic panel | Challenge pooling and fitted-unit assumptions | Supported contracts are respected and exclusions are explicit |
| Missing or failed evidence | Challenge release discipline | The affected use is blocked rather than silently passed |
| Reloaded and altered scenario | Challenge reproducibility | Replay works and a changed basis cannot reuse the original review |
Use common supplied data and experimental evidence across comparable entries. A synthetic case can establish behaviour under its declared generating process; it does not establish validity for a client. Real data without credible causal evidence cannot supply causal ground truth. A product-specific case is allowed for an additional capability, but must be labelled outside the common comparison.
Review record
For each case, retain a completed record with the following fields. This is a team-maintained review record, not a new AMMM3 YAML configuration schema.
| Record field | Required content |
|---|---|
| Identity | Protocol version, case ID, date, product and exact version or commit |
| Scope | Estimator, intended use, estimand and explicitly excluded claims |
| Frozen inputs | Case specification, data/configuration hashes and environment reference |
| Assessment rows | M1 through M8, applicability, evidence state, native/extension/manual mechanism and evidence paths |
| Evidence basis | Expected behaviour, actual observation, failure-case result and reviewer rationale for each applicable requirement |
| Interpretation | Numerical, statistical and causal limits; completed sensitivity findings |
| Decision | Accepted for a stated use, restricted to a narrower use, or blocked |
| Accountability | Modeller, technical reviewer, business owner and any lack of independent review |
| Follow-up | Outstanding work, expiry or review triggers, superseded record and amendments |
Missing mandatory evidence produces a blocked decision for the proposed use. A reviewer must not fill the gap with confidence in the vendor, an LLM summary or an undocumented recollection. Where reviewers disagree, retain the disputed requirement and evidence. The claim remains blocked or restricted until the technical disagreement is resolved in a signed record.
Fair treatment of the comparison set
The benchmark concerns demonstrated behaviour within a declared workflow. It must not assume that Bayesian methods are causally identified or that regularised methods lack all uncertainty procedures.
PyMC-Marketing documents configurable Bayesian models, calibration, optimisation and production workflow components. Include those components when they form the assessed workflow, and record any required integration. PyMC-Marketing guidance provides the starting point; it is not benchmark evidence of execution.
Meridian documents model health checks and comparative fit/holdout measures. The absence of a preferred information criterion would not alone fail M5. Assess the evidence relevant to the intended use. Meridian health checks provide the initial capability reference.
Robyn documents regularised model fitting, calibration and bootstrap intervals within clusters of candidate models. Evaluate what those intervals condition on and whether they support the claimed estimand; do not equate them with a joint Bayesian posterior or describe uncertainty as wholly absent. Robyn features provide the initial capability reference.
These primary references were checked on 11 September 2026. Record exact versions and sources again for each assessment. A changing documentation page cannot serve as a frozen execution record. Commercial workflows must meet the same evidence rules; inaccessible capabilities remain not assessed.
Publication boundary
A public capability matrix may report evidence states, supported scope, mechanisms and linked review records. It must distinguish an implemented feature from a demonstrated workflow. Unsupported operations and unresolved limitations must remain visible for AMMM3 as well as competitors.
Do not publish runtime comparisons, numerical performance rankings, aggregate scores or claims that this protocol establishes a superior causal estimator. The acceptable conclusion is that a named workflow demonstrated specified requirements for a stated use, subject to its recorded limits.
Amendment record
| Version | Date | Change |
|---|---|---|
| 1.0 | 11 September 2026 | Freeze the methodological requirements, case families, evidence states and review procedure; establish no comparative results |