Estimator Capabilities

Current implementation status

This document is the approved Gate A contract for hardening the random-effects (RE), correlated-random-effects (CRE), and fixed-effects (FE) estimators. Approval fixes the intended scope. It does not mean that every target capability is implemented or production qualified.

The current implementation status is authoritative until the relevant implementation and evidence gates have passed:

Estimator Current status Safe interpretation
RE supported_with_limits Existing Gaussian hierarchical estimator. Its population slopes require conditional mean independence between group effects and regressors. Retained block-support diagnostics are available; Monte Carlo qualification is pending.
CRE supported_with_limits Enforced random-intercept Mundlak augmentation for simple externally transformed or untransformed numeric regressors. Retained between-design and hierarchy block-support diagnostics are available; Monte Carlo qualification is pending.
FE implemented_not_qualified The direct R API and bounded YAML runner implement coefficient-only Gaussian FE with exact within-unit contrasts. P4-J technical gates passed, but deterministic one-replication checks do not provide repeated-sampling qualification. FE does not recover unit intercepts or provide level prediction.

No estimator is production qualified by this document alone. Production status requires the implementation, package validation, Stan review where applicable, and predeclared Monte Carlo evidence.

Shared interpretation boundary

Estimator choice changes how time-invariant group heterogeneity is handled. It does not establish causal identification. RE, CRE, and FE can all be biased by omitted time-varying confounding, reverse causality, measurement error, or a misspecified response function. Applied work must assess those risks separately.

The package will fail closed for a structurally unsupported estimator and design combination. Advisory diagnostics will remain distinct from structural errors and will report their metric, threshold, and recovery action.

Estimator design-diagnostic policy

diagnostics_policy_thresholds() owns one nested estimator_design policy. This keeps estimator-specific design checks separate from the existing general design-matrix and fitted-model diagnostics. In particular, the estimator condition-number rule does not replace the existing kappa_* policy, and the within-share rule does not replace the existing within_var_* policy.

The version 1 estimator-design policy fixes these boundaries:

Check Pass boundary Advisory boundary Structural or strict failure
SVD rank tolerance Full column rank at max(n, p) * max(s) * .Machine$double.eps Not applicable Numerical rank below the column count at that tolerance
Zero within norm Above .Machine$double.eps * max(1, max(abs(x))) * sqrt(length(x)) Not applicable At or below the resulting tolerance
Within share >= 0.05 < 0.05 Exact-zero within variation remains structural
Variance inflation factor (VIF) <= 20 > 20 Non-finite results remain structural
Centred and scaled condition number <= 30 > 30 Non-finite results remain structural
Between-design residual degrees of freedom >= 2 < 2 in explore and publish modes < 2 in strict mode
Group levels >= 4 2 or 3 Fewer than 2

The group-level boundary is an operational small-support warning, not a claim that four groups guarantee reliable variance-component estimation. The VIF, condition-number, within-share, residual-degrees-of-freedom, and group-support boundaries are conservative screening rules rather than universal statistical constants.

For YAML runner fits, diagnostics.policy_mode is retained before fitting and selects this estimator-design policy during CRE pre-flight and preparation. Consequently, strict mode rejects a full-rank combined between-design with fewer than two residual degrees of freedom before Stan compilation. Direct API preparation uses the publish policy.

P2-A defines and tests this policy. P2-B through P2-E implement the numerical reports, apply the policy to the exact retained frame, and retain the resulting evidence. The Stan model and estimation algebra are unchanged. Structurally invalid retained designs now fail before fitting; finite support limitations remain advisory. These diagnostics do not qualify an estimator for production use.

Capability matrix

The following matrix records the implemented FE v1 boundary and the approved RE and CRE scope. Implementation does not imply production qualification.

Capability RE target CRE v1 target FE v1 target
Group structure One or more existing hierarchy blocks One selected CRE group within an otherwise valid hierarchy Exactly one unit identifier
Panel keys Existing hierarchy/date contract Existing hierarchy/date contract One unit key and one date key; retained unit-date pairs must be unique
Group effects Existing Gaussian random intercepts and slopes A random intercept is required in the selected group Unit intercepts are removed by exact within contrasts
Random slopes Existing support Rejected in the selected CRE group; other hierarchy blocks retain RE semantics Rejected
Balanced panels Supported Supported Supported
Unbalanced panels Supported Supported Supported when every retained unit has at least two dates
Internally transformed media Existing support Rejected when a CRE variable overlaps an internally transformed channel Rejected
External numeric transforms Supported Supported as simple one-column main effects Supported as simple one-column main effects
Factors and interactions Existing formula contract Allowed only when they do not involve CRE variables Rejected in v1
Scaling Existing global contract Existing contract after retained-sample CRE augmentation Existing global contract followed by within contrasts
Offsets Existing hierarchy contract Same as RE Rejected in v1
MCMC Supported Supported Supported
MAP Existing support Existing support Rejected in v1
Level fitted values and prediction Existing support Existing support with the retained CRE lookup Rejected in v1
New-group prediction Existing hierarchy contract Existing contract plus a fitted CRE lookup Rejected
Time-series cross-validation Existing supported cases Supported only with fold-local CRE means Rejected in v1
LOO, WAIC, and run comparison Existing supported cases Existing supported cases Rejected because contrast rows are not original-observation likelihood terms
Decomposition and optimisation Existing supported cases Existing support; CRE mean terms are not causal media contributions Rejected in v1

Random effects

Target estimator

RE remains the existing Gaussian hierarchical estimator. Group coefficients are partially pooled through the declared random-effects blocks. This can be a useful variance model when its conditional independence assumptions are credible.

Retained support diagnostics

DSAMbayes retains diagnostics for each random-effects block:

  • group count and group-size distribution;
  • random-effects design dimensions and rank by group;
  • all-zero and non-finite random-effect columns;
  • within-group variation for random slopes; and
  • covariance dimension relative to available groups.

All-zero columns, non-finite matrices, fewer than two group levels, and invalid dimensions are structural errors. The response is also rejected as a random-effect design term or grouping key. Weak within variation, per-group rank loss, and small group counts are advisory. Covariance dimension and the groups-per-covariance-parameter ratio are reported without an identification threshold. A group-constant random slope is not automatically described as mathematically unidentified when between-group information may still inform its covariance.

Assumption boundary

Plain RE does not protect population slopes when the random intercept is correlated with included regressors. In that case, use CRE only if its narrower conditional-mean specification is credible, or use another design.

Correlated random effects

Target estimator

CRE v1 is a random-intercept Mundlak specification. For each selected numeric regressor, DSAMbayes adds its mean over the exact retained estimation rows for the selected group. The coefficient on the original regressor is the within slope. The coefficient on its group mean is the contextual difference. Their sum is the between slope.

Supported scope

CRE v1 requires:

  • one explicitly selected hierarchy group;
  • a random intercept and no random slope in that selected group;
  • a CRE variable that maps to one simple numeric population-design column;
  • no overlap between that variable and an enabled internal media transform;
  • group means computed from the same retained rows and fitted covariate basis used by Stan; and
  • holdout, cross-validation, and deployment lookups derived from the fitted training frame only.

Other hierarchy blocks may retain their ordinary RE semantics. Repeated selected-group/date pairs are allowed when another hierarchy level provides multiple observations in the same period.

Rejected specifications and recovery

Rejected condition Reason Recovery
Random slope in the selected CRE group Additive means do not correct correlated random slopes Remove that random slope from the selected group or use RE and state its assumption
CRE variable uses an internal adstock or Hill transform The fitted regressor changes with posterior parameters, so a raw-column mean is not the fitted-basis mean Precompute and declare a fixed transformed numeric column, or omit CRE for that channel
CRE variable appears in a factor, interaction, spline, polynomial, or inline transform The variable does not map to one unambiguous fitted column Author a simple numeric column explicitly
Selected group has no random intercept The approved Mundlak contract concerns correlation with a random intercept Add a random intercept or do not request CRE
Generated mean name collides with a user column Silent replacement would make the fitted basis ambiguous Rename the user column or choose a non-conflicting CRE prefix
New group is absent from the fitted lookup Its training group mean is undefined Fit with the group represented or use an estimator with an approved new-group contract

Retained-sample rule

set_cre() will become configuration-only. At fit time, DSAMbayes will build one explicit completeness and finiteness mask across the response, population terms, random-effects terms, grouping keys, date, offset, and transformation inputs. CRE means, default priors, design diagnostics, Stan data, and retained lookups will all derive from that frozen frame. na.rm = TRUE will not define the fitted group mean.

CRE does not correct omitted time-varying confounding or guarantee that the conditional mean of the group effect is linear in the included means.

Fixed effects

Implemented estimator

FE v1 is a separate fixed_effects model class constructed with fixed_effects(formula, data, unit). Before fitting, the caller must supply exactly one date mapping through the existing set_date() lifecycle.

Surface Status Boundary
Direct R API fitting Implemented MCMC coefficient and residual-noise inference within the FE v1 contract
Runner configuration Implemented Use model.type: fe plus fixed_effects.unit; validation and dry-run do not compile or sample
Runner fitting and artefacts Implemented, not qualified Non-dry runs use the existing MCMC fit and dedicated coefficient, sampler, within-design, contrast-residual, and contrast posterior-predictive artefacts
Qualification Technical gates passed; production qualification not complete P4-J bounded runner, prior-only, one-replication recovery, and sampler gates passed; P5 repeated Monte Carlo remains separately gated

The non-dry YAML path uses the same estimator, prior, boundary, scaling, and sampling contracts as the direct R API. It returns before generic level-scale diagnostics, model selection, scenario analysis, optimisation, forecasting, or deployment. A completed runner result has no publishability claim, and its factual FE diagnostics summary records qualification_status: not_assessed.

For each retained unit, FE builds a deterministic orthonormal Helmert contrast matrix Q_i and fits:

Q_i y_i ~ Normal(Q_i X_i beta, sigma)

This removes unrestricted unit intercepts without estimating shrinkage priors for them. Unequal unit sizes are allowed. Singleton units, duplicate retained unit-date keys, non-finite values, time-invariant regressors, and rank-deficient within designs will abort before sampling.

Version 1 output contract

FE v1 supports slope and noise inference, within-design diagnostics, and contrast-space posterior predictive checks. Its strongest possible release label is production for coefficient inference only.

FE v1 rejects:

  • level-scale fitted() and predict() output;
  • conditional unit-intercept reconstruction;
  • unseen-unit prediction;
  • decomposition and counterfactual response;
  • budget optimisation;
  • time-series cross-validation;
  • pointwise LOO, WAIC, and run comparison;
  • MAP fitting; and
  • internal media transformations.

Contrast-space fitted values are diagnostics. They are not level-scale MMM fitted values or predictions. Adding level outputs later requires a separate posterior contract, linear oracle, coverage study, and counterfactual specification.

The P4-G bounded synthetic integration fit proves only that the direct API can compile, sample, extract its declared output, and survive a same-environment RDS round-trip. P4-I proves the bounded runner orchestration and artefact contract with a mocked Stan boundary. P4-J adds bounded live runner and explicit-prior prior-only checks plus balanced and unbalanced coefficient, noise, and sampler regression gates. All passed their predeclared technical checks.

The two P4-J recovery fits are deterministic one-replication regression evidence. They do not estimate repeated-sampling coverage or production reliability. They also do not estimate bias, root mean squared error, failure rates, or robustness. FE therefore remains implemented_not_qualified. Repeated Monte Carlo evidence and its separate review remain P5 work.

Qualification requirements

Each estimator receives its own release decision. A package-wide pass cannot hide a failed estimator cell.

Qualification requires:

  1. deterministic design and algebra oracles;
  2. typed structural failures before Stan compilation;
  3. four-chain sampler diagnostics on every declared estimand and likelihood scale parameter;
  4. repeated paired Monte Carlo evidence under RE-valid, RE-invalid/CRE-valid, unbalanced, weak-design, and misspecified scenarios;
  5. conditional coverage and unconditional success-and-cover rates, so failed fits remain visible;
  6. Monte Carlo uncertainty for bias, coverage, and paired loss comparisons;
  7. bounded prior-sensitivity evidence; and
  8. human statistical and Stan review where model code changes.

Possible release labels are production, experimental, and unsupported. Any failed core gate blocks production. An inconclusive stress cell narrows the supported capability and remains visible in the evidence report.