CRE / Mundlak
Purpose
The correlated random effects (CRE) pathway, implemented as a Mundlak device, augments hierarchical DSAMbayes models with group-mean terms. This separates within-group variation from between-group variation for selected regressors. It addresses correlation between group effects and regressors when the conditional mean of the selected random intercept is adequately represented by the included group means. It does not correct correlated random slopes, time-varying endogeneity, or omitted time-varying confounders.
When to use CRE
Use CRE when:
- The model is hierarchical (panel data with
(term | group)syntax). - Time-varying regressors (e.g. externally transformed media signals) have group-level means that may be correlated with the selected group intercept.
- You want to decompose effects into within-group (temporal) and between-group (cross-sectional) components.
Do not use CRE when:
- The model is BLM or pooled (CRE requires hierarchical class).
- The panel has only one group (no between-group variation exists).
- All regressors of interest are time-invariant (CRE mean terms would be constant).
- The selected group needs random slopes, or a CRE variable uses an internal draw-dependent adstock or Hill transform.
Construction
CRE is applied after model construction via set_cre():
What set_cre() does
-
Resolves the grouping variable. If the formula has one group factor, it is used automatically. If multiple group factors exist, the
groupargument must be specified explicitly. -
Generates group-mean column names. For each variable in
vars, a mean-term column is namedcre_mean_<variable>(configurable viaprefix). -
Preserves the source data.
set_cre()does not add generated columns to.original_data. It creates a provisional prior view for the generated terms; fit preparation resolves the final priors from the retained frame. -
Updates the formula. The generated mean terms are appended to the population formula as fixed effects.
-
Builds authoritative means at fit time. DSAMbayes reconciles all population, random-effect, group, date, offset, and transform dependencies; excludes incomplete or non-finite rows; then calculates group means over exactly the rows and fitted covariate basis passed to Stan.
-
Extends priors and boundaries. Default prior and boundary entries are added for each new mean term. Explicit prior overrides remain unchanged.
YAML runner configuration
When using the runner, CRE is configured via:
The runner calls set_cre() during model construction for model.type: cre.
Mundlak decomposition
For a regressor $x_{gt}$ (group $g$, time $t$), DSAMbayes fits the equivalent parameterisation
$$ y_{gt} = \beta x_{gt} + \theta \bar{x}_g + \ldots $$where:
- Within-group effect: $\beta$, the coefficient on $x_{gt}$ after conditioning on the group mean.
- Contextual difference: $\theta$, the coefficient on $\bar{x}_g$. It is the difference between the between-group and within-group slopes.
- Total between-group slope: $\beta + \theta$.
The original coefficient on $x_{gt}$ in a standard random-effects model conflates both sources. Adding $\bar{x}_g$ as a fixed effect separates them.
Validation and identification warnings
Input validation
set_cre() validates:
- The model is hierarchical (aborts for BLM or pooled).
- The selected group exists and contains a random intercept but no random slope.
- Each CRE variable is one simple numeric population main effect. Factors, interactions, inline transforms, polynomials, and splines involving a CRE variable are rejected.
- No CRE variable overlaps an enabled internal draw-dependent media transform.
- No generated CRE mean term appears in a random-effects block.
- Generated column names do not collide with user-authored source columns.
The same typed structural checks run in direct data preparation, pre-flight,
fit(), fit_map(), authored v2 config validation, compiled v1 validation,
and runner preparation. pre_flight_strict = FALSE cannot downgrade them.
Identification warnings
warn_cre_identification() checks two conditions after CRE setup:
-
More CRE variables than groups. If
length(vars) > n_groups, between-effect estimates may be weakly identified. The function emits a warning. -
Near-zero within-group variation. For each CRE variable, the within-group residual ($x_{gt} - \bar{x}_g$) standard deviation is checked. If it is effectively zero, within-effect identification is weak. The function emits a per-variable warning.
Zero-variance CRE mean terms
If the combined between-group design is rank-deficient, pre-flight reports the
problem and strict runner fitting aborts. Reduce hierarchy.cre_variables, add
credible group variation, or use RE only when its independence assumption is
defensible. Disabling scaling does not repair a rank-deficient CRE design.
Panel assumptions
- Balanced panels are not required. Each retained observation contributes once to its group mean, so unequal group sizes are supported.
- Missing and non-finite dependencies are reconciled jointly. A row omitted
from the likelihood cannot influence a fitted CRE mean. Mean calculation does
not use
na.rm = TRUEafter the retained frame is frozen. - Generated columns are owned by DSAMbayes. Fit preparation rebuilds them from source variables. A user-authored column with a reserved generated name aborts instead of being overwritten.
- Prediction uses fitted means. Holdout, cross-validation, and deployment paths use training-only lookups derived from the fitted estimation frame.
Estimator evidence smoke
The default suite runs deterministic checks of the repository-owned RE/CRE DGP
and evidence aggregation contracts. An explicit Stan smoke checks that one
fixed RI-02 replication can be fitted through the public RE and CRE APIs:
This smoke asserts only that usable posterior coefficient draws exist. It does
not measure recovery, calibration, bias, relative estimator performance, or
production readiness. The immutable manifest, terminal-record, and aggregation
workflow under scripts/estimator-evidence/ owns pilot and confirmatory Monte
Carlo evidence. Do not treat a smoke result as estimator qualification or
causal validation.
Retained diagnostics evidence
Runner fits retain two CRE-specific diagnostic artefacts under
40_diagnostics/:
estimator_checks.csvrecords the typed structural-contract result and the pre-flight CRE configuration result;row_reconciliation.csvrecords input, retained, and excluded counts plus excluded source-row IDs and reason codes. It does not contain response values.
These artefacts establish which checks and rows the fit used. They do not establish causal identification or Monte Carlo calibration.
Fixed-effects boundary
DSAMbayes does not currently implement the planned exact within-unit fixed-effects (FE) estimator. A hierarchical random-intercept model is not an FE estimator. Do not label it as FE or use it as a substitute for the pending implementation and evidence gates.
Decomposition and reporting
CRE mean terms appear as ordinary fixed-effect terms in the population formula. This means:
- Posterior summary includes CRE mean-term coefficients alongside other population coefficients.
- Response decomposition via
decomp()attributes fitted-value contributions to CRE mean terms separately from their within-group counterparts. - Plots (posterior forest, prior-vs-posterior) include CRE mean terms.
Interpretation note: the CRE mean-term coefficient is a contextual difference, not the total between-group slope. For a model containing $x_{gt}$ and $\bar{x}_g$, calculate the total between-group slope as the sum of their two coefficients. Treat all three quantities as associational unless the research design supports a causal interpretation.
Cross-references
- Model Classes, hierarchical class construction
- Priors and Boundaries, prior schema for CRE terms
- Config Schema,
model.type: creandhierarchy.*YAML keys - Diagnostics Gates, within-group variation check