Diagnostics Gates
For the workflow interpretation of these checks, start with Stage 4: Computation and Sampler and Stage 5: Model Adequacy. This page is the threshold and policy reference.
Use this page when you need exact gate thresholds, status aggregation, or YAML policy semantics. It does not replace substantive model review: a model can clear threshold tables and still be a poor basis for interpretation.
Model selection for time-ordered data
PSIS-LOO and WAIC treat pointwise observations as conditionally exchangeable. That assumption is not generally appropriate for time-ordered MMM data, where nearby weeks can remain dependent after conditioning on the fitted model.
Use the runner’s expanding-window blocked CV or leave-future-out CV as the primary evidence when selecting among time-series MMM specifications. Treat PSIS-LOO, WAIC, and Pareto-k outputs as supplementary fit and influence diagnostics. They do not establish future-period predictive performance, causal validity, or a publish-gate pass on their own.
Purpose
DSAMbayes runs a deterministic diagnostics framework after model fitting. Each diagnostic check produces a pass, warn, or fail status. The policy mode controls how lenient or strict the thresholds are. This page defines the check taxonomy, threshold tables, policy modes, identifiability gate, and the overall status aggregation rule.
How to use this page
- Use Stage 4: Computation and Sampler to understand which checks are non-negotiable before trusting the posterior.
- Use this page to see the exact DSAMbayes thresholds and artifact semantics.
- Use Stage 5: Model Adequacy before treating a passing diagnostics table as permission for decomposition, model comparison, or optimisation.
Policy modes
The diagnostics framework supports three policy modes, configured via diagnostics.policy_mode in YAML:
| Mode | Intent | Threshold behaviour |
|---|---|---|
explore |
Rapid iteration during model development | Relaxed fail thresholds; many checks can only warn, not fail |
publish |
Default production mode for shareable outputs | P0 sampler and integrity checks can fail; condition-number, residual, boundary and within-variation P1 checks are warn-only |
strict |
Audit-grade gating for release candidates | Tightest thresholds; rank deficit fails rather than warns |
The mode is resolved by diagnostics_policy_thresholds(mode) in R/diagnostics_report.R.
Check taxonomy
Checks are organised into phases:
| Phase | Scope | When evaluated |
|---|---|---|
P0 |
Data integrity and critical MCMC reliability | Pre-fit and post-fit publish gates |
P1 |
Design conditioning, residual behaviour and identifiability | Pre-fit and post-fit model-review checks |
P2 |
Supplementary model-selection evidence | Post-fit predictive scoring |
Each check row includes:
| Field | Meaning |
|---|---|
check_id |
Unique identifier |
phase |
P0, P1 or P2 |
severity |
Check priority recorded by the diagnostics registry |
status |
pass, warn, fail, or skipped |
metric |
Metric name |
value |
Observed value |
threshold |
Applied threshold description |
message |
Human-readable explanation |
Design checks
| Check ID | Metric | Pass | Warn | Fail |
|---|---|---|---|---|
pre_response_finite |
non_finite_response_count |
== 0 |
n/a | > 0 |
pre_design_constants_duplicates |
constant_plus_duplicate_columns |
== 0 |
n/a | > 0 |
pre_design_rank_deficit |
rank_deficit |
== 0 |
> 0 (publish) |
> 0 (strict) |
pre_design_condition_number (P1) |
kappa_X |
≤ warn |
> warn |
> fail |
Condition number thresholds by mode
| Mode | Warn | Fail |
|---|---|---|
explore |
10,000 | ∞ (cannot fail) |
publish |
10,000 | 1,000,000 (warn-only; fail disabled) |
strict |
10,000 | 1,000,000 |
P0 sampler checks (MCMC only)
| Check ID | Metric | Direction | Warn | Fail |
|---|---|---|---|---|
post_rhat_max |
max_rhat |
Lower is better | 1.01 | 1.01 |
post_ess_bulk_min |
min_ess_bulk |
Higher is better | 400 | 200 |
post_ess_tail_min |
min_ess_tail |
Higher is better | 200 | 100 |
post_ebfmi |
min_ebfmi |
Higher is better | 0.30 | 0.20 |
post_treedepth_saturation |
treedepth_hit_fraction |
Lower is better | 0.00 | 0.01 |
post_divergences |
divergent_fraction |
Lower is better | 0.00 | 0.00 |
Mode adjustments for sampler checks
DSAMbayes treats any Rhat above 1.01 as a failure in publish and strict
modes, following the rank-normalised, folded Rhat guidance in
Vehtari et al. (2021) and the
Stan warnings guide. In
explore mode, the fail threshold is deliberately relaxed to 1.10, while
the warning threshold remains 1.01.
P1 residual checks
| Check ID | Metric | Direction | Warn | Fail |
|---|---|---|---|---|
post_residual_ljung_box_p_min |
resid_lb_p |
Higher is better | 0.05 | 0.01 |
post_residual_max_abs_acf |
resid_acf_max |
Lower is better | 0.20 | 0.40 |
Mode adjustments for residual checks
| Mode | resid_lb_p warn |
Raw fail boundary / effective state | resid_acf warn |
Raw fail boundary / effective state |
|---|---|---|---|---|
explore |
0.05 | 0.00 (cannot fail) | 0.20 | ∞ (cannot fail) |
publish |
0.05 | 0.01 (warn-only; fail disabled) | 0.20 | 0.40 (warn-only; fail disabled) |
strict |
0.10 | 0.05 | 0.15 | 0.30 |
In publish mode, values beyond the raw residual fail boundary remain warn.
The persisted threshold field states that failure is disabled and retains the
raw boundary for triage. Non-residual P0 failures can still make the overall
run fail.
P1 boundary hit check
| Check ID | Metric | Direction | Warn | Fail |
|---|---|---|---|---|
post_boundary_hit_rate_max |
boundary_hit_frac |
Lower is better | 0.05 | 0.20 |
In explore mode, boundary hits cannot fail. Publish mode retains the raw
fail > 0.20 boundary for triage but records values beyond it as warn.
Failure is disabled for this check in publish mode. In strict mode,
thresholds tighten to warn > 0.02, fail > 0.10.
P1 within-group variation check
| Check ID | Metric | Direction | Warn | Fail |
|---|---|---|---|---|
pre_within_variation_ratio_min |
within_var_min_ratio |
Higher is better | 0.10 | 0.05 |
This check applies to hierarchical models and flags groups where within-group
variation is extremely low relative to between-group variation. In explore
mode, the fail threshold is zero (cannot fail). Publish mode retains the raw
fail < 0.05 boundary for triage but records values below it as warn;
failure is disabled for this check. Strict mode applies warn < 0.15 and
fail < 0.10.
Identifiability gate
The P1 pre_identifiability_baseline_media_corr check measures the maximum
absolute correlation between baseline terms and media terms in the design
matrix. It is configured via diagnostics.identifiability in YAML:
Term detection
- Media terms: explicitly listed in
media_terms. - Baseline terms: union of
baseline_terms, generated time-component terms, and matches frombaseline_regexpatterns. - Both sets are intersected with actual design-matrix columns and filtered to remove constant columns.
Thresholds by mode
| Mode | Warn | Fail |
|---|---|---|
explore |
0.80 | ∞ (cannot fail) |
publish |
0.80 | 0.95 |
strict |
0.70 | 0.85 |
Skip conditions
The identifiability gate reports skipped when:
identifiability.enabled: false- No configured media terms found in the design matrix
- No baseline terms detected from configured terms/regex
- All resolved baseline or media terms are constant
Overall status aggregation
The overall diagnostics status is determined by diagnostics_overall_status():
- If any check has
status == "fail"→ overall status is fail. - If any check has
status == "warn"(and none fail) → overall status is warn. - Otherwise → overall status is pass.
Checks with status == "skipped" do not affect the overall status.
P2 rows use the emitted post_model_selection_method, post_psis_loo,
post_psis_loo_options_ignored, post_psis_loo_time_dependence,
post_psis_loo_elpd, post_psis_loo_looic and
post_psis_loo_pareto_k identifiers. They record supplementary predictive
evidence and do not override P0 or P1 failures.
The runner also emits post_sampler_params_available when sampler parameters
can be inspected and post_sampler_unavailable_for_optimise for the explicit
MAP not-applicable path. These identifiers are part of the persisted report;
they are not substitute MCMC diagnostics for a MAP fit.
Runner artefact output
The diagnostics framework produces:
| Artefact | Location | Content |
|---|---|---|
diagnostics_report.csv |
40_diagnostics/ |
Full check table with all fields |
diagnostics_summary.txt |
40_diagnostics/ |
Human-readable summary of overall status and failing checks |
Interpretation guidance
pass: no enabled check breached its configured threshold. This does not establish substantive adequacy, causal validity or suitability for the intended decision.warn: review recommended; the model may have quality concerns but does not block the configured policy.fail: at least one enabled check breached a fail threshold; resolve it before using the result for decisions.
Common remediation actions
| Diagnostic area | Warning signs | Actions |
|---|---|---|
| High Rhat | > 1.01 |
Increase MCMC iterations or warmup; simplify model |
| Low ESS | < 400 bulk or < 200 tail |
Increase iterations; check for multimodality |
| Divergences | Any non-zero fraction | Increase adapt_delta; reparameterise model |
| High condition number | kappa > 10,000 |
Reduce collinearity; remove redundant terms |
| Residual autocorrelation | High ACF or low Ljung-Box p | Add time controls (trend, seasonality, holidays) |
| Boundary hits | > 5% of draws |
Review boundary specification; widen or remove constraints |
| High baseline-media correlation | > 0.80 |
Add controls to separate baseline from media; consider alternative model specifications |
Cross-references
- Model Classes, which diagnostics apply to each class
- Response Scale Semantics, residuals are computed on model scale
- Config Schema,
diagnostics.*YAML keys - Output Artefacts, diagnostics artefact paths
- Diagnostics Plots, visual diagnostic outputs