Prevent a negative baseline

In Meridian, the baseline represents the expected outcome under the counterfactual scenario where all treatment variables are set to their baseline values (zero for paid and organic media; the configured baseline value for non-media treatments). Because the outcome can't be negative, a negative baseline indicates statistical error in the causal inference of the treatment effects, typically meaning the model is over-crediting the treatments. Some statistical error is expected in any model, so an occasional, small dip below zero isn't cause for concern, but an extremely negative baseline (or a high posterior probability that the baseline aggregated over the entire time window is negative) is a clear signal that the model needs adjustment.

Setting allows_negative_aggregate_baseline=False in ModelSpec results in a negative baseline probability of 0.0 as evaluated by the Negative Baseline model health check. It constrains the aggregate outcome baseline to be non-negative on every prior and posterior draw, where aggregate means summed across all geos and time periods (while baseline values for individual geos or time periods can still be negative).

By default, allows_negative_aggregate_baseline=True. To enable the constraint, set allows_negative_aggregate_baseline=False:

from meridian.model import model
from meridian.model import spec

model_spec = spec.ModelSpec(allows_negative_aggregate_baseline=False)
mmm = model.Meridian(input_data=data, model_spec=model_spec)

How the constraint works

The expected baseline is the sum of several model components: time effects (knot_values; see Set knots), geo main effects (tau_g_excl_baseline), control variables, and non-media treatments evaluated at their baseline values.

When allows_negative_aggregate_baseline=False, Meridian reparameterizes knot_values using an orthogonal change of basis (a Householder reflection) so that the aggregate baseline across all geos and time periods is governed by a single parameter, while the remaining \(K - 1\) components remain unconstrained to capture time effects.

During prior and posterior sampling, Meridian samples the geo, control, and non-media treatment parameters first, and then places a TruncatedNormal prior on that single knot parameter with a dynamic lower bound that depends on those sampled components, ensuring a non-negative aggregate outcome baseline on every draw. Baseline values for individual geos or time periods can still be negative when warranted by the data.

For full technical details, see the Mathematical appendix.

Key properties

Constraining the aggregate baseline to be non-negative provides several desirable statistical and computational properties:

  • Zero probability of a negative aggregate baseline: Every prior and posterior draw satisfies the non-negativity constraint on the aggregate outcome baseline, ensuring that Analyzer.negative_baseline_probability() equals 0.0 when evaluated over all geos and time periods and that the Negative Baseline component in ModelReviewer returns PASS (0.0 probability). Evaluating a subset of geos or time periods (selected_geos, selected_times), or the KPI baseline (use_kpi=True) when revenue_per_kpi varies across geos or time periods, is not constrained and can return a non-zero probability.
  • Preserves causal deconfounding: Parameters intended to debias treatment effects—including geo main effects, control variable coefficients, and the remaining \(K - 1\) components of knot_values—are not truncated or constrained.
  • Compatible with all model structures without tuning: Turned on with a single boolean switch (allows_negative_aggregate_baseline=False) without requiring penalty tuning or changes to Meridian's default priors, and works seamlessly across both geo-level and national-level models, any knot specification (a single constant intercept \(K = 1\), manual knot lists, full time effects \(K = T\), or Automatic Knot Selection), control variables, non-media treatments, organic media, and reach and frequency (RF) channels.
  • No MCMC sampling slowdown: Because the orthogonal change-of-basis matrix is precomputed in closed form and only a single parameter is truncated, the posterior geometry remains smooth for the No-U-Turn Sampler (NUTS) with negligible computational overhead.
  • Minimal impact when the baseline is already positive: Truncating a single transformed knot parameter is a minimally sufficient prior adjustment. When the data and priors already support a positive aggregate baseline, the truncation threshold lies far in the left tail of the distribution, so posterior estimates change only slightly compared to the unconstrained model. Because the setting changes the prior, results still differ for every model, and the same random seed produces different draws under the two settings.

Interaction with priors and Model Health

Constraining the aggregate baseline to be non-negative interacts with your model priors and can surface prior-data conflicts in Model Health diagnostics.

Treatment priors and Model Health diagnostics

When the non-negative baseline constraint is enabled, you may occasionally see other Model Health components impacted if your treatment priors conflict with a non-negative aggregate baseline.

A negative baseline in an unconstrained model may arise from overly optimistic treatment priors. When priors on ROI, mROI, or contribution imply a total treatment contribution that exceeds the observed total outcome, constraining the aggregate baseline to be non-negative creates a conflict between the priors and the data, which can cause the posterior predictive total outcome to exceed the observed total outcome.

When this conflict occurs, the constraint prevents the model from hiding excess incremental outcome inside a negative baseline. Instead, the conflict can surface in ModelReviewer diagnostics:

If you observe convergence issues, a Bayesian PPP failure, or lower R-squared after enabling the constraint, inspect and recalibrate your treatment priors rather than allowing a negative baseline. For step-by-step guidance, see Mitigate negative or low baseline.

Effect on the knot_values prior and custom prior requirements

Under the constraint, the induced prior on knot_values is no longer independent and identically distributed (i.i.d.) Normal. Because the first transformed knot coordinate is truncated at a lower bound that depends on the geo, control, and non-media treatment parameters, the reconstructed knot_values are correlated, shifted upward in aggregate, and dependent on those parameters. The marginal priors on the geo, control, and non-media treatment parameters themselves are not modified, because they are sampled before the truncated knot coordinate.

Custom priors on all other model parameters—such as media roi_m, mroi_m, contribution_m, tau_g_excl_baseline, gamma_c, and gamma_n—are unrestricted and sampled directly from their specified distributions.

You can inspect the resulting prior with sample_prior. The prior draws contain the reconstructed knot_values; the internal parameters rotated_knot_0 and rotated_knot_rest are not saved.

from meridian.analysis import analyzer
from meridian.model import model
from meridian.model import spec

model_spec = spec.ModelSpec(allows_negative_aggregate_baseline=False)
mmm = model.Meridian(input_data=data, model_spec=model_spec)
mmm.sample_prior(500)

# Prior draws of knot_values with dimensions (chain, draw, knots).
knot_values_prior = mmm.inference_data.prior.knot_values

# Prior probability of a negative aggregate baseline; 0.0 under the constraint.
analyzer.Analyzer(mmm).negative_baseline_probability(use_posterior=False)

Comparing these draws with those from an unconstrained model (allows_negative_aggregate_baseline=True) shows how the constraint changes the knot_values prior. After running sample_posterior, you can also plot the prior and posterior of knot_values with ModelDiagnostics.plot_prior_and_posterior_distribution(parameter='knot_values').

Mathematical appendix

This section describes the mathematical formulation of the non-negative aggregate baseline constraint. For additional details on the mathematical notation used in this section, see Mathematical notation and Input data.

1. Aggregate baseline condition on the transformed KPI scale

Let \(y_{g,t} = (\ddot{y}_{g,t}/p_g - m^{[Y]}) / s^{[Y]}\) denote the centered and population-scaled KPI. Under the counterfactual baseline scenario where all paid and organic media variables are set to zero and non-media treatments are set to their baseline values \(x^{[N](0)}_{g,t,i}\), the expected transformed KPI at geo \(g\) and time \(t\) is:

\[ \hat{y}^{\text{base}}_{g,t} := E\left(Y_{g,t} \;\middle|\; \left\{x^{(0)}_{g,t,i}\right\}, \left\{z_{g,t,i}\right\}\right) = \mu_t + \tau_g + \sum\limits_{i=1}^{N_C} \gamma^{[C]}_{g,i} z_{g,t,i} + \sum\limits_{i=1}^{N_N} \gamma^{[N]}_{g,i} x^{[N](0)}_{g,t,i} \]

where:

  • \(\mu_t = \sum_{k=1}^{K} W_{t,k} b_k\) is the time-varying intercept determined by the \(T \times K\) knot weight matrix \(W\) and knot parameters \(b = (b_1, \dots, b_K)^\top\) (see How the knots argument works).
  • \(\tau_g\) is the geo main effect.
  • \(\gamma^{[C]}_{g,i} z_{g,t,i}\) is the contribution of control variable \(i\).
  • \(\gamma^{[N]}_{g,i} x^{[N](0)}_{g,t,i}\) is the contribution of non-media treatment variable \(i\) evaluated at its counterfactual baseline value \(x^{[N](0)}_{g,t,i}\).

Let \(u^{[Y]}_{g,t}\) denote revenue_per_kpi (with \(u^{[Y]}_{g,t} = 1\) when revenue_per_kpi is None), and define normalized outcome weights \(w_{g,t}\):

\[ w_{g,t} = \frac{T \cdot p_g u^{[Y]}_{g,t}}{\sum\limits_{g'=1}^{G} \sum\limits_{t'=1}^{T} p_{g'} u^{[Y]}_{g',t'}} \]

so that \(\sum_{g=1}^{G} \sum_{t=1}^{T} w_{g,t} = T\). Inverting the KPI scaling transformation shows that the aggregate baseline outcome summed over all geos and time periods is non-negative if and only if:

\[ \sum\limits_{k=1}^{K} W^*_k b_k + \text{offset}\left(\tau, \gamma^{[C]}, \gamma^{[N]}\right) \ge -\frac{T \cdot m^{[Y]}}{s^{[Y]}} \]

where \(W^*_k = \sum_{t=1}^{T} W_{t,k} \left(\sum_{g=1}^{G} w_{g,t}\right)\) is the total outcome-weighted contribution of knot \(k\), and \(\text{offset}(\tau, \gamma^{[C]}, \gamma^{[N]})\) aggregates the outcome-weighted geo, control, and baseline non-media treatment terms:

\[ \text{offset}\left(\tau, \gamma^{[C]}, \gamma^{[N]}\right) = \sum\limits_{g=1}^{G} \sum\limits_{t=1}^{T} w_{g,t} \left( \tau_g + \sum\limits_{i=1}^{N_C} \gamma^{[C]}_{g,i} z_{g,t,i} + \sum\limits_{i=1}^{N_N} \gamma^{[N]}_{g,i} x^{[N](0)}_{g,t,i} \right) \]

2. Orthogonal change of basis for knot_values

To isolate the linear combination \((W^*)^\top b = \sum_{k=1}^{K} W^*_k b_k\) into a single coordinate without altering the prior geometry, Meridian constructs a \(K \times K\) symmetric orthogonal matrix \(Q\) (\(Q Q^\top = I_K\)) as a closed-form Householder reflection whose first row is the unit vector \(u = W^* / \|W^*\|_2\).

Defining the transformed knot parameters \(\tilde{b} = (\tilde{b}_1, \tilde{b}_2, \dots, \tilde{b}_K)^\top = Q b\) (so that \(b = Q^\top \tilde{b}\); in the model computation graph, \(\tilde{b}_1\) is named rotated_knot_0 and \((\tilde{b}_2, \dots, \tilde{b}_K)\) is named rotated_knot_rest):

  • The aggregate knot contribution depends exclusively on the first transformed coordinate \(\tilde{b}_1\):

    \[ (W^*)^\top b = (W^*)^\top Q^\top \tilde{b} = \|W^*\|_2 \, \tilde{b}_1 \]

  • Because \(Q\) is orthogonal, the isotropic zero-mean Normal prior \(b_k \stackrel{\text{i.i.d.}}{\sim} \text{Normal}(0, \sigma_b)\) is invariant under orthogonal transformations and induces mutually independent \(\text{Normal}(0, \sigma_b)\) priors on the transformed coordinates \(\tilde{b}_1, \tilde{b}_2, \dots, \tilde{b}_K\):

    \[ b_k \stackrel{\text{i.i.d.}}{\sim} \text{Normal}(0, \sigma_b) \quad \text{for } k = 1, \dots, K \quad \Longleftrightarrow \quad \tilde{b}_k \stackrel{\text{i.i.d.}}{\sim} \text{Normal}(0, \sigma_b) \quad \text{for } k = 1, \dots, K \]

3. Dynamic truncation on \(\tilde{b}_1\)

During prior and posterior sampling, Meridian samples \(\tau_g\), \(\gamma^{[C]}_{g,i}\), and \(\gamma^{[N]}_{g,i}\) first, evaluates the dynamic lower bound:

\[ \tilde{b}_{1,\text{low}}\left(\tau, \gamma^{[C]}, \gamma^{[N]}\right) = \frac{-\frac{T \cdot m^{[Y]}}{s^{[Y]}} - \text{offset}\left(\tau, \gamma^{[C]}, \gamma^{[N]}\right)}{\|W^*\|_2} \]

and places a TruncatedNormal prior on the single parameter \(\tilde{b}_1\):

\[ \tilde{b}_1 \sim \text{TruncatedNormal}\left(\text{loc}=0, \text{scale}=\sigma_b, \text{low}=\tilde{b}_{1,\text{low}}, \text{high}=\tilde{b}_{1,\text{high}}\right) \]

where the upper truncation bound is set to:

\[ \tilde{b}_{1,\text{high}} = \max\left(\tilde{b}_{1,\text{low}} + 50\sigma_b, \; 50\sigma_b\right) \]

The upper bound is a numerical requirement, not a modeling constraint. The No-U-Turn Sampler (NUTS) samples \(\tilde{b}_1\) on an unconstrained scale through the sigmoid transformation that TensorFlow Probability's TruncatedNormal uses, which requires finite bounds. Setting the upper bound at least \(50\sigma_b\) greater than both zero and \(\tilde{b}_{1,\text{low}}\) leaves negligible prior probability in the upper tail, so it behaves as \(+\infty\) in practice.

The remaining \(K - 1\) transformed knot parameters \(\tilde{b}_2, \dots, \tilde{b}_K \stackrel{\text{i.i.d.}}{\sim} \text{Normal}(0, \sigma_b)\) remain unconstrained Normal variables, and Meridian reconstructs the original knot parameters as \(b = Q^\top \tilde{b}\).