In Meridian, the baseline represents the expected outcome under the counterfactual scenario where all treatment variables are set to their baseline values (zero for paid and organic media; the configured baseline value for non-media treatments). Because the outcome can't be negative, a negative baseline indicates statistical error in the causal inference of the treatment effects, typically meaning the model is over-crediting the treatments. Some statistical error is expected in any model, so an occasional, small dip below zero isn't cause for concern, but an extremely negative baseline (or a high posterior probability that the baseline aggregated over the entire time window is negative) is a clear signal that the model needs adjustment.
Setting allows_negative_aggregate_baseline=False in
ModelSpec results in
a negative baseline
probability
of 0.0 as evaluated by the Negative
Baseline model
health check. It constrains the aggregate outcome baseline to be non-negative on
every prior and posterior draw, where aggregate means summed across all geos
and time periods (while baseline values for individual geos or time periods can
still be negative).
By default, allows_negative_aggregate_baseline=True. To enable the constraint,
set allows_negative_aggregate_baseline=False:
from meridian.model import model
from meridian.model import spec
model_spec = spec.ModelSpec(allows_negative_aggregate_baseline=False)
mmm = model.Meridian(input_data=data, model_spec=model_spec)
How the constraint works
The expected baseline is the sum of several model components: time effects
(knot_values; see Set knots),
geo main effects (tau_g_excl_baseline), control variables, and non-media
treatments evaluated at their baseline values.
When allows_negative_aggregate_baseline=False, Meridian reparameterizes
knot_values using an orthogonal change of basis (a Householder reflection) so
that the aggregate baseline across all geos and time periods is governed by a
single parameter, while the remaining \(K - 1\) components remain
unconstrained to capture time effects.
During prior and posterior sampling, Meridian samples the geo, control, and
non-media treatment parameters first, and then places a TruncatedNormal prior
on that single knot parameter with a dynamic lower bound that depends on those
sampled components, ensuring a non-negative aggregate outcome baseline on every
draw. Baseline values for individual geos or time periods can still be negative
when warranted by the data.
For full technical details, see the Mathematical appendix.
Key properties
Constraining the aggregate baseline to be non-negative provides several desirable statistical and computational properties:
- Zero probability of a negative aggregate baseline: Every prior and
posterior draw satisfies the non-negativity constraint on the aggregate
outcome baseline, ensuring that
Analyzer.negative_baseline_probability()equals0.0when evaluated over all geos and time periods and that the Negative Baseline component inModelReviewerreturnsPASS(0.0probability). Evaluating a subset of geos or time periods (selected_geos,selected_times), or the KPI baseline (use_kpi=True) whenrevenue_per_kpivaries across geos or time periods, is not constrained and can return a non-zero probability. - Preserves causal deconfounding: Parameters intended to debias treatment
effects—including geo main effects, control variable coefficients, and the
remaining \(K - 1\) components of
knot_values—are not truncated or constrained. - Compatible with all model structures without tuning: Turned on with a
single boolean switch (
allows_negative_aggregate_baseline=False) without requiring penalty tuning or changes to Meridian's default priors, and works seamlessly across both geo-level and national-level models, any knot specification (a single constant intercept \(K = 1\), manual knot lists, full time effects \(K = T\), or Automatic Knot Selection), control variables, non-media treatments, organic media, and reach and frequency (RF) channels. - No MCMC sampling slowdown: Because the orthogonal change-of-basis matrix is precomputed in closed form and only a single parameter is truncated, the posterior geometry remains smooth for the No-U-Turn Sampler (NUTS) with negligible computational overhead.
- Minimal impact when the baseline is already positive: Truncating a single transformed knot parameter is a minimally sufficient prior adjustment. When the data and priors already support a positive aggregate baseline, the truncation threshold lies far in the left tail of the distribution, so posterior estimates change only slightly compared to the unconstrained model. Because the setting changes the prior, results still differ for every model, and the same random seed produces different draws under the two settings.
Interaction with priors and Model Health
Constraining the aggregate baseline to be non-negative interacts with your model priors and can surface prior-data conflicts in Model Health diagnostics.
Treatment priors and Model Health diagnostics
When the non-negative baseline constraint is enabled, you may occasionally see other Model Health components impacted if your treatment priors conflict with a non-negative aggregate baseline.
A negative baseline in an unconstrained model may arise from overly optimistic treatment priors. When priors on ROI, mROI, or contribution imply a total treatment contribution that exceeds the observed total outcome, constraining the aggregate baseline to be non-negative creates a conflict between the priors and the data, which can cause the posterior predictive total outcome to exceed the observed total outcome.
When this conflict occurs, the constraint prevents the model from hiding excess
incremental outcome inside a negative baseline. Instead, the conflict can
surface in ModelReviewer
diagnostics:
- Convergence:
Elevated
R-hatvalues (max_r_hat >= 1.2) if strong treatment priors push the sampler against the non-negative baseline boundary. - Bayesian Posterior Predictive P-value
(PPP): A
FAILresult (p-value< 0.05) when the posterior predictive total outcome exceeds the observed total outcome. - Goodness-of-fit: A decrease in R-squared.
If you observe convergence issues, a Bayesian PPP failure, or lower R-squared after enabling the constraint, inspect and recalibrate your treatment priors rather than allowing a negative baseline. For step-by-step guidance, see Mitigate negative or low baseline.
Effect on the knot_values prior and custom prior requirements
Under the constraint, the induced prior on knot_values is no longer
independent and identically distributed (i.i.d.) Normal. Because the first
transformed knot coordinate is truncated at a lower bound that depends on the
geo, control, and non-media treatment parameters, the reconstructed
knot_values are correlated, shifted upward in aggregate, and dependent on
those parameters. The marginal priors on the geo, control, and non-media
treatment parameters themselves are not modified, because they are sampled
before the truncated knot coordinate.
Custom priors on all other model parameters—such as media roi_m, mroi_m,
contribution_m, tau_g_excl_baseline, gamma_c, and gamma_n—are
unrestricted and sampled directly from their specified distributions.
You can inspect the resulting prior with sample_prior. The prior draws
contain the reconstructed knot_values; the internal parameters
rotated_knot_0 and rotated_knot_rest are not saved.
from meridian.analysis import analyzer
from meridian.model import model
from meridian.model import spec
model_spec = spec.ModelSpec(allows_negative_aggregate_baseline=False)
mmm = model.Meridian(input_data=data, model_spec=model_spec)
mmm.sample_prior(500)
# Prior draws of knot_values with dimensions (chain, draw, knots).
knot_values_prior = mmm.inference_data.prior.knot_values
# Prior probability of a negative aggregate baseline; 0.0 under the constraint.
analyzer.Analyzer(mmm).negative_baseline_probability(use_posterior=False)
Comparing these draws with those from an unconstrained model
(allows_negative_aggregate_baseline=True) shows how the constraint changes the
knot_values prior. After running sample_posterior, you can also plot the
prior and posterior of knot_values with
ModelDiagnostics.plot_prior_and_posterior_distribution(parameter='knot_values').
Mathematical appendix
This section describes the mathematical formulation of the non-negative aggregate baseline constraint. For additional details on the mathematical notation used in this section, see Mathematical notation and Input data.
1. Aggregate baseline condition on the transformed KPI scale
Let \(y_{g,t} = (\ddot{y}_{g,t}/p_g - m^{[Y]}) / s^{[Y]}\) denote the centered and population-scaled KPI. Under the counterfactual baseline scenario where all paid and organic media variables are set to zero and non-media treatments are set to their baseline values \(x^{[N](0)}_{g,t,i}\), the expected transformed KPI at geo \(g\) and time \(t\) is:
\[ \hat{y}^{\text{base}}_{g,t} := E\left(Y_{g,t} \;\middle|\; \left\{x^{(0)}_{g,t,i}\right\}, \left\{z_{g,t,i}\right\}\right) = \mu_t + \tau_g + \sum\limits_{i=1}^{N_C} \gamma^{[C]}_{g,i} z_{g,t,i} + \sum\limits_{i=1}^{N_N} \gamma^{[N]}_{g,i} x^{[N](0)}_{g,t,i} \]
where:
- \(\mu_t = \sum_{k=1}^{K} W_{t,k} b_k\) is the time-varying intercept
determined by the \(T \times K\) knot weight matrix \(W\) and knot
parameters \(b = (b_1, \dots, b_K)^\top\) (see How the
knotsargument works). - \(\tau_g\) is the geo main effect.
- \(\gamma^{[C]}_{g,i} z_{g,t,i}\) is the contribution of control variable \(i\).
- \(\gamma^{[N]}_{g,i} x^{[N](0)}_{g,t,i}\) is the contribution of non-media treatment variable \(i\) evaluated at its counterfactual baseline value \(x^{[N](0)}_{g,t,i}\).
Let \(u^{[Y]}_{g,t}\) denote revenue_per_kpi (with \(u^{[Y]}_{g,t} = 1\)
when revenue_per_kpi is None), and define normalized outcome weights
\(w_{g,t}\):
\[ w_{g,t} = \frac{T \cdot p_g u^{[Y]}_{g,t}}{\sum\limits_{g'=1}^{G} \sum\limits_{t'=1}^{T} p_{g'} u^{[Y]}_{g',t'}} \]
so that \(\sum_{g=1}^{G} \sum_{t=1}^{T} w_{g,t} = T\). Inverting the KPI scaling transformation shows that the aggregate baseline outcome summed over all geos and time periods is non-negative if and only if:
\[ \sum\limits_{k=1}^{K} W^*_k b_k + \text{offset}\left(\tau, \gamma^{[C]}, \gamma^{[N]}\right) \ge -\frac{T \cdot m^{[Y]}}{s^{[Y]}} \]
where \(W^*_k = \sum_{t=1}^{T} W_{t,k} \left(\sum_{g=1}^{G} w_{g,t}\right)\) is the total outcome-weighted contribution of knot \(k\), and \(\text{offset}(\tau, \gamma^{[C]}, \gamma^{[N]})\) aggregates the outcome-weighted geo, control, and baseline non-media treatment terms:
\[ \text{offset}\left(\tau, \gamma^{[C]}, \gamma^{[N]}\right) = \sum\limits_{g=1}^{G} \sum\limits_{t=1}^{T} w_{g,t} \left( \tau_g + \sum\limits_{i=1}^{N_C} \gamma^{[C]}_{g,i} z_{g,t,i} + \sum\limits_{i=1}^{N_N} \gamma^{[N]}_{g,i} x^{[N](0)}_{g,t,i} \right) \]
2. Orthogonal change of basis for knot_values
To isolate the linear combination \((W^*)^\top b = \sum_{k=1}^{K} W^*_k b_k\) into a single coordinate without altering the prior geometry, Meridian constructs a \(K \times K\) symmetric orthogonal matrix \(Q\) (\(Q Q^\top = I_K\)) as a closed-form Householder reflection whose first row is the unit vector \(u = W^* / \|W^*\|_2\).
Defining the transformed knot parameters
\(\tilde{b} = (\tilde{b}_1, \tilde{b}_2, \dots, \tilde{b}_K)^\top = Q b\)
(so that \(b = Q^\top \tilde{b}\); in the model computation graph,
\(\tilde{b}_1\) is named rotated_knot_0 and
\((\tilde{b}_2, \dots, \tilde{b}_K)\) is named rotated_knot_rest):
The aggregate knot contribution depends exclusively on the first transformed coordinate \(\tilde{b}_1\):
\[ (W^*)^\top b = (W^*)^\top Q^\top \tilde{b} = \|W^*\|_2 \, \tilde{b}_1 \]
Because \(Q\) is orthogonal, the isotropic zero-mean Normal prior \(b_k \stackrel{\text{i.i.d.}}{\sim} \text{Normal}(0, \sigma_b)\) is invariant under orthogonal transformations and induces mutually independent \(\text{Normal}(0, \sigma_b)\) priors on the transformed coordinates \(\tilde{b}_1, \tilde{b}_2, \dots, \tilde{b}_K\):
\[ b_k \stackrel{\text{i.i.d.}}{\sim} \text{Normal}(0, \sigma_b) \quad \text{for } k = 1, \dots, K \quad \Longleftrightarrow \quad \tilde{b}_k \stackrel{\text{i.i.d.}}{\sim} \text{Normal}(0, \sigma_b) \quad \text{for } k = 1, \dots, K \]
3. Dynamic truncation on \(\tilde{b}_1\)
During prior and posterior sampling, Meridian samples \(\tau_g\), \(\gamma^{[C]}_{g,i}\), and \(\gamma^{[N]}_{g,i}\) first, evaluates the dynamic lower bound:
\[ \tilde{b}_{1,\text{low}}\left(\tau, \gamma^{[C]}, \gamma^{[N]}\right) = \frac{-\frac{T \cdot m^{[Y]}}{s^{[Y]}} - \text{offset}\left(\tau, \gamma^{[C]}, \gamma^{[N]}\right)}{\|W^*\|_2} \]
and places a TruncatedNormal prior on the single parameter \(\tilde{b}_1\):
\[ \tilde{b}_1 \sim \text{TruncatedNormal}\left(\text{loc}=0, \text{scale}=\sigma_b, \text{low}=\tilde{b}_{1,\text{low}}, \text{high}=\tilde{b}_{1,\text{high}}\right) \]
where the upper truncation bound is set to:
\[ \tilde{b}_{1,\text{high}} = \max\left(\tilde{b}_{1,\text{low}} + 50\sigma_b, \; 50\sigma_b\right) \]
The upper bound is a numerical requirement, not a modeling constraint. The
No-U-Turn Sampler (NUTS) samples \(\tilde{b}_1\) on an unconstrained scale
through the sigmoid transformation that TensorFlow Probability's
TruncatedNormal uses, which requires finite bounds. Setting the upper bound at
least \(50\sigma_b\) greater than both zero and
\(\tilde{b}_{1,\text{low}}\) leaves negligible prior probability in the upper
tail, so it behaves as \(+\infty\) in practice.
The remaining \(K - 1\) transformed knot parameters \(\tilde{b}_2, \dots, \tilde{b}_K \stackrel{\text{i.i.d.}}{\sim} \text{Normal}(0, \sigma_b)\) remain unconstrained Normal variables, and Meridian reconstructs the original knot parameters as \(b = Q^\top \tilde{b}\).