Model Classes

Purpose

DSAMbayes provides four model classes for Bayesian marketing mix modelling. Each class targets a different data structure and pooling strategy. This page describes the constructor pathways, fit support, and practical limitations of each class so that an operator can select the appropriate model for a given dataset.

The fixed-effects (FE) class is a separate coefficient-only estimator. A hierarchical model with unit random intercepts remains a random-effects model, not an FE estimator.

Use this page to choose the right modelling surface for your data structure and decision problem. Do not use it as a substitute for the broader workflow: class selection does not settle prior design, computational trustworthiness, model adequacy, or causal interpretation. For that framing, start with What Principled Means, Stage 2: Model and Priors, and Stage 5: Model Adequacy.

Selection discipline

  • Choose the simplest class that matches the real data structure.
  • Do not move to hierarchical or pooled models only because they look more advanced.
  • Treat model class as a structural choice, not proof that the resulting model is decision-ready.
  • After choosing a class, return to the workflow pages for prior-setting, diagnostics, and interpretation discipline.

Class summary

Class S3 class chain Constructor Data structure Grouping Typical use case
BLM blm blm(formula, data) Single market/brand None One-market regression with full prior and boundary control
Hierarchical hierarchical, blm blm(formula, data) with (term | group) syntax Panel (long format) Random effects by group Multi-market models sharing strength across groups
Pooled pooled, blm pool(blm_obj, grouping_vars, map) Single market Structured coefficient pooling via dimension map Single-market models with media coefficients pooled across labelled dimensions
Fixed effects fixed_effects, requires_prior, blm fixed_effects(formula, data, unit) then set_date() Panel (long format) Exact within-unit contrasts; common slopes Within-unit coefficient inference when unrestricted unit intercepts are nuisance parameters

Fixed effects (fixed_effects)

Construction

model <- fixed_effects(kpi ~ m_tv + m_search + trend, panel_df, market) |>
  set_date(week)

FE v1 removes unrestricted unit intercepts with deterministic orthonormal within-unit contrasts. It estimates common slopes and residual noise through MCMC. Formula terms must be bare numeric columns, and retained unit-date keys must be unique. Unequal unit sizes and irregular date spacing are allowed.

Output boundary

Use get_posterior() for reporting-scale slope and residual-noise draws and within-contrast generated quantities. FE v1 does not recover unit intercepts or support MAP, level fitted values, prediction, model selection, decomposition, counterfactual analysis, optimisation, forecasting, or deployment. It supports YAML validation, dry-run, and bounded MCMC fitting with dedicated coefficient and contrast-space artefacts.

P4-J’s bounded live runner, explicit-prior prior-only, one-replication recovery, and sampler gates passed. These checks do not estimate repeated-sampling coverage or production reliability. The implementation is not production-qualified, so FE remains implemented_not_qualified. See Estimator Capabilities for the full contract.

BLM (blm)

Construction

model <- blm(kpi ~ m_tv + m_search + trend + seasonality, data = df)

blm() dispatches on the first argument. When passed a formula, it creates a blm object with default priors and boundaries. When passed an lm object, it creates a bayes_lm_updater whose priors are initialised from the OLS coefficient estimates and standard errors.

Fit support

Method Function Backend
MCMC fit(model, ...) rstan::sampling()
MAP fit_map(model, n_runs, ...) rstan::optimizing() (repeated starts)

Post-fit accessors

  • get_posterior(), coefficient draws, fitted values, metrics
  • fitted(), predicted response on the original scale
  • decomp(), native predictor-level decomposition. It returns a dsambayes_decomposition S3 result for BLM objects; use $DecompedData for the long-form contribution table. Hierarchical models return a keyed tibble with one such result in each row’s decomp list-column.
  • optimise_budget(), decision-layer budget allocation

Decomposition migration

decomp() no longer returns the external DSAMdecomp S4 calculators. This is a breaking development-version change. Replace S4 slot access such as result@DecompedData with result$DecompedData; use result$DecompGroupProps for grouped totals. The DSAMdecomp-only plotting and post-processing helpers are not part of the native result. Build any required downstream summaries from these two tables.

Limitations

  • No group structure. For multi-market data, use the hierarchical class.
  • optimise_budget() aborts if scale=TRUE and an offset is present (unsupported combination for the bayes_lm_updater Stan template).

Hierarchical (hierarchical)

Construction

The hierarchical class is created automatically when blm() detects random-effects syntax (|) in the formula:

model <- blm(
  kpi ~ m_tv + m_search + trend + (1 + m_tv + m_search | market),
  data = panel_df
)

Terms to the left of | become random slopes; the variable to the right defines the grouping factor. Multiple grouping terms are supported.

The ordinary random-effects specification partially pools group effects and uses both within-group and between-group information for population coefficients. Its interpretation requires the group effects to be conditionally mean-independent of the regressors. If this assumption is not credible, use a justified CRE specification or another design. Neither option solves omitted time-varying confounding by itself.

CRE / Mundlak extension

For correlated random effects, call set_cre() after construction:

cre_model <- blm(
  kpi ~ m_tv_signal + m_search_signal + trend + (1 | market),
  data = panel_df
) |>
  set_cre(vars = c("m_tv_signal", "m_search_signal"))

CRE v1 is restricted to a selected random-intercept group with no random slope in that group. Each CRE variable must be one simple numeric main effect and must not overlap an internal draw-dependent media transform. Fit preparation calculates the Mundlak means from the exact retained estimation frame. See CRE / Mundlak for details.

Fit support

Method Function Backend
MCMC fit(model, ...) rstan::sampling()
MAP fit_map(model, n_runs, ...) rstan::optimizing() (repeated starts)

Post-fit accessors

Same as BLM. Coefficient draws from get_posterior() return vectors (one value per group) rather than scalars. Budget optimisation uses the population-level (fixed-effect) coefficient draws from the beta parameter.

Limitations

  • Stan template compilation uses a templated source (general_hierarchical.stan) rendered per number of groups and parameterisation mode. First compilation is slow; subsequent runs use a cached binary.
  • Response decomposition is native and returns one result per retained group. Probabilistic media transforms remain unsupported for decomposition and are rejected explicitly.
  • Formula offsets enter native decomposition as coefficient-one baseline components and must be included in the decomposition mapping table.
  • Posterior forest and prior-vs-posterior plots average group-specific draws to produce a single population-level estimate.
  • Offset support in the hierarchical Stan template is handled via stats::model.offset() within build_hierarchical_frame_data().

Pooled (pooled)

Construction

The pooled class is created by converting an existing BLM object with pool():

base <- blm(kpi ~ m_tv + m_search + trend + seasonality, data = df)
model <- pool(base, grouping_vars = c("channel"), map = pooling_map)

The map is a data frame with a variable column mapping formula terms to pooling dimension labels. Exact formula-term labels are preferred; raw variable names are accepted only when they resolve unambiguously to a single non-offset formula term. Priors and boundaries are reset to defaults when pool() is called.

Fit support

Method Function Backend
MCMC fit(model, ...) rstan::sampling()

MAP fitting (fit_map) is not currently implemented for pooled models.

Post-fit accessors

Same as BLM. The design matrix is split into base terms (intercept + non-pooled) and media terms (pooled). The Stan template uses a per-dimension coefficient structure.

Limitations

  • MAP fitting is not available.
  • extract_stan_design_matrix() may return a zero-row matrix, which causes VIF computation to be skipped.
  • The pooled Stan cache key includes sorted grouping variable names to avoid collisions between different pooling configurations.
  • Time-series cross-validation is available for pooled MCMC models, subject to the same media-transform restrictions as other classes.

Class selection guide

Scenario Recommended class Rationale
Single market, sufficient data BLM Simplest pathway; full accessor and optimisation support
Single market, OLS baseline available BLM via blm(lm_obj, data) Priors initialised from OLS; Bayesian updating
Multi-market panel Hierarchical Partial pooling shares strength across markets
Multi-market panel where the selected random intercept may correlate with regressors Hierarchical + CRE Mundlak terms model that correlation through included group means when the random-intercept conditional-mean specification is adequate
Single market with structured media dimensions Pooled Coefficient pooling across labelled media categories

In practice, the class decision should usually be driven by three questions:

  1. Is the dataset a single time series or a grouped panel?
  2. Do you need partial pooling across real groups, or pooling across labelled coefficient dimensions?
  3. Is the added structure necessary for the business question, or are you adding complexity without a clear identifiability benefit?

Fit method selection

Criterion MCMC (fit) MAP (fit_map)
Full posterior Yes No (point estimate only)
Credible intervals Yes No; restart diagnostics only
Diagnostics (Rhat, ESS, divergences) Yes Not applicable
LOO-CV / model selection Yes Not supported
Speed Minutes to hours Seconds to minutes
Budget optimisation Full posterior-based Point-estimate-based

For inferential runs where diagnostics and uncertainty quantification matter, MCMC is the recommended fit method. MAP is useful for rapid iteration during model development. Fit method alone does not qualify an estimator for production use.

MAP returns one selected optimum, not posterior draws. Do not derive credible intervals or MCMC diagnostics from it; its point estimate can understate uncertainty, especially for hierarchical variance components. fit_map() retains the restart results for inspection (and runner outputs include optimisation_runs.csv), so materially different restart objectives should be treated as optimisation instability or competing local optima, not as a substitute for posterior uncertainty.

Cross-references