Getting Started

Purpose

Onboard a new user from install to first successful DSAMbayes run, then point them into the workflow guidance needed for serious modelling use.

Audience

  • New DSAMbayes users.
  • Analysts running DSAMbayes through R scripts or CLI.
  1. Install and Setup, or Install and Setup (Windows / RStudio) on an admin-restricted Windows machine
  2. Concepts
  3. Your First BLM Model or Your First Hierarchical Model
  4. Quickstart (YAML Runner)
  5. FAQ
  6. Principled Bayesian Workflow

If you are coming from classical econometrics, read Frequentist to Bayesian Translation before customising priors or interpreting diagnostics.

Pages

Page Topic
Install and Setup Prerequisites, installation commands, and verification
Install and Setup (Windows / RStudio) Setup guidance for admin-restricted Windows machines running RStudio
Concepts What DSAMbayes does and how Bayesian MMM differs from classical regression
Your First BLM Model Build, fit, and interpret a single-market model using the R API
Your First Hierarchical Model Multi-market model with partial pooling and CRE
Quickstart (YAML Runner) Minimal end-to-end CLI run from config to output inspection
FAQ Answers to common questions

Use Getting Started for setup, not for the whole methodology

The Getting Started pages are intentionally practical. Once you can run the package, use the workflow section to answer the two questions that matter most in Bayesian MMM:

  • how should I think about priors?
  • which diagnostics matter before I trust outputs?

Subsections of Getting Started

Install and Setup

Audience

Engineers and analysts setting up DSAMbayes for local development or modelling runs. This page is written for a Unix shell (Linux or macOS). If you are on Windows using RStudio, especially on a work-issued machine without full administrator rights, use Install and Setup (Windows / RStudio) instead.

Prerequisites

  • R >= 4.1: check with R --version.
  • A C++ toolchain for Stan compilation. This is the most common source of setup issues:
    • macOS: install Xcode Command Line Tools (xcode-select --install).
    • Windows: install Rtools matching your R version. Ensure make is on your PATH.
    • Linux (Ubuntu/Debian): sudo apt install build-essential.
    • See the RStan Getting Started Guide for detailed platform instructions.
  • A local checkout of this repository.

Open a terminal in the repository root and run:

# 1. Select an ABI-safe host library and create it with the cache
source scripts/r-library-path.sh
dsambayes_set_r_library host
mkdir -p "$R_LIBS_USER" .cache

# 2. Set the cache path (add both settings to .bashrc/.zshrc for persistence)
export XDG_CACHE_HOME="$PWD/.cache"

# 3. Install DSAMbayes from the local checkout
R -q -e 'install.packages(".", repos = NULL, type = "source")'

This keeps all package libraries and Stan compilation caches inside the repo, avoiding permission issues with system library paths.

Do not share a compiled package library between R runtimes or containers. dsambayes_set_r_library host selects .Rlib-host-r<active-R-version>/; use dsambayes_set_r_library container inside a container instead. The path is deliberately derived from the active runtime, not hard-coded in the repository.

Verify the installation

1. Confirm DSAMbayes loads

R -q -e 'library(DSAMbayes); cat("Version:", as.character(utils::packageVersion("DSAMbayes")), "\n")'

Expected: prints the installed package version.

2. Confirm the runner works

Rscript scripts/dsambayes.R validate --config config/blm_timeseries.yaml

Expected: validation completes without errors.

3. (Optional) Run the test suite

R -q -e 'testthat::test_dir("tests/testthat")'

Expected: all tests pass.

Alternative: install the private fork from GitHub

If you need only the package API and have read access to the private canonical repository, install it directly:

R -q -e 'remotes::install_github("tandpds/DSAMbayes-Charles-Dev", dependencies = NA, upgrade = "never")'

This installs the current package version from main. remotes reads GITHUB_PAT from the environment for private-repository access. Keep the token in your user environment or credential manager, never in a command, script, or tracked file. Clone the repository instead when you need its example configs, data, runme.R, or development tooling.

Using renv (optional)

The repository includes a renv.lock file for fully reproducible dependency management. To use it:

R -q -e 'install.packages("renv"); renv::restore()'

This installs the exact dependency versions used during development. It is optional but recommended for production runs where reproducibility matters.

Runner setup and first execution

1. Validate the example config

Rscript scripts/dsambayes.R validate --config config/blm_timeseries.yaml

2. Execute a full run

Rscript scripts/dsambayes.R run --config config/cre_geo_panel.yaml

Expected: a timestamped run directory is created under results/ with model outputs and diagnostics.

Troubleshooting

Stan compilation fails

Symptom: errors during Compiling model... referencing C++ or compiler issues.

Actions:

  1. Confirm your C++ toolchain is working: R -q -e 'pkgbuild::has_build_tools(debug = TRUE)'.
  2. On Windows, ensure Rtools is installed and make is on your PATH.
  3. Clear the Stan cache and retry: rm -rf .cache/dsambayes.
  4. Follow the RStan Getting Started Guide for your platform.

Package installation fails

Symptom: install.packages(".", repos = NULL, type = "source") errors.

Actions:

  1. Confirm you are in the repository root directory.
  2. Confirm the selected R_LIBS_USER directory exists and is writable.
  3. Check for missing system dependencies in the error output.

Stale Stan cache

Symptom: unexpected model behaviour after updating the package.

Actions:

  1. Clear the cache: rm -rf .cache/dsambayes.
  2. Re-run with model.force_recompile: true in your config if you need to invalidate a stale compiled model.

Permission issues

Symptom: write failures for library, cache, or run outputs.

Actions:

  1. Ensure R_LIBS_USER, .cache, and results/ are writable.
  2. Keep R_LIBS_USER and XDG_CACHE_HOME set in your shell session.
  3. Run all commands from the repository root.

Concepts

Purpose

Give new DSAMbayes users a compact conceptual orientation before they move into tutorials, runner usage, or the workflow section.

This page is intentionally introductory. It does not try to be the full methodology guide for Bayesian MMM. For that, use Principled Bayesian Workflow.

What is DSAMbayes?

DSAMbayes is an R package for Bayesian marketing mix modelling built on Stan. It provides:

  • an lm()-style modelling interface for interactive work
  • model classes for single-series, hierarchical, and pooled MMM
  • prior and boundary controls
  • post-fit extraction, diagnostics, decomposition, and optimisation tooling
  • a YAML/CLI runner for reproducible runs

The main practical difference from classical regression is that DSAMbayes works with a posterior distribution, not just a single fitted coefficient vector.

Why Bayesian MMM?

MMM datasets often have the exact features that make naive regression unstable:

  • short time series
  • overlapping media timing
  • strong baseline structure
  • uncertain functional form
  • real business need for uncertainty-aware decisions

Bayesian modelling helps because it makes several things explicit:

  • regularisation through priors
  • structural constraints through boundaries
  • uncertainty propagation into downstream outputs
  • diagnostic gates rather than fit-statistic-only thinking

The DSAMbayes mental model

DSAMbayes should be thought of as a workflow, not just a fitter.

At a high level:

  1. specify the model and priors
  2. fit the model
  3. check whether the posterior computation is trustworthy
  4. check whether the fitted model is adequate for the data
  5. only then interpret decomposition, comparison, or optimisation outputs

That is the main philosophical shift from a simpler OLS-style workflow.

Model classes

DSAMbayes supports three main model classes.

BLM

Single-market Bayesian linear model.

Use when:

  • you have one KPI series
  • one market / brand / region is the modelling unit
  • you want the simplest Bayesian MMM starting point

Hierarchical

Multi-group model with partial pooling.

Use when:

  • you have panel data across markets, regions, or brands
  • you want to borrow strength across groups while preserving group structure

Pooled

Single-market model with structured pooling across labelled media dimensions.

Use when:

  • the outcome is one series
  • the media structure has nested or repeated dimensions that should share information

For the detailed class contract, see Model Classes.

Interactive API vs runner

DSAMbayes has two main ways of working:

Interactive R API

Best when you want to prototype directly in R:

  • blm()
  • set_prior()
  • set_boundary()
  • fit()
  • get_posterior()

YAML / CLI runner

Best when you want reproducibility and staged artefacts:

  • validate
  • run
  • staged outputs under results/

See Quickstart and Runner.

What this page does not try to teach

This page does not try to fully answer:

  • how to choose priors
  • which diagnostics matter most
  • when business interpretation is allowed

Those are workflow questions, and they are handled in the dedicated methodology pages:

Your First BLM Model

Goal

Build, fit, and interpret a single-market Bayesian linear model (BLM) using the DSAMbayes R API.

This page is a hands-on tutorial. For the broader methodological questions behind prior-setting and diagnostics, use the workflow pages:

Prerequisites

  • DSAMbayes installed locally (see Install and Setup).
  • Familiarity with R and lm()-style formulas.

Dataset

This walkthrough uses the tracked synthetic dataset at data/timeseries/demo_data_synthetic.csv. It contains weekly observations for a single series with columns for:

  • Response: revenue, weekly synthetic KPI.
  • Media: channel0_spend, channel1_spend, channel2_spend, channel3_spend.
  • Controls: events, newsletters.
  • Date: time, weekly date index.
library(DSAMbayes)
df <- read.csv("data/timeseries/demo_data_synthetic.csv")
str(df)

Step 1: Construct the model

blm() creates an unfitted model object. No fitting happens yet.

model <- blm(
  revenue ~
    events + newsletters +
    channel0_spend + channel1_spend + channel2_spend + channel3_spend,
  data = df
) %>%
  set_date(time)

Inspect defaults:

peek_prior(model)      # data-dependent normal priors: normal(ybar, sy) intercept, normal(0, sy / sx) per slope
peek_boundary(model)   # (-Inf, Inf), unconstrained

Step 2: Set boundaries

Media channels should have non-negative effects. Use inequality notation:

model <- model %>%
  set_boundary(
    channel0_spend > 0, channel1_spend > 0,
    channel2_spend > 0, channel3_spend > 0
  )

Step 3: (Optional) Override priors

Default priors are weakly informative. Override only with domain knowledge:

model <- model %>%
  set_prior(events ~ normal(0, 1))

See Stage 2: Model and Priors and Minimal-Prior Policy for guidance.

Step 4: Fit with MCMC

fitted_model <- model %>%
  fit(cores = 2, iter = 2000, warmup = 1000, seed = 123)

First-time Stan compilation takes 1–3 minutes. Subsequent runs use a cached binary. With the default 4 chains split across 2 cores, sampling on synthetic data typically completes in under 2 minutes.

Step 5: Sampler diagnostics

chain_diagnostics(fitted_model)
Metric Good Concern
Max Rhat <= 1.01 > 1.01 means chains have not converged
Min ESS (bulk) > 400 < 200 means too few effective samples
Divergences 0 Any non-zero count warrants investigation

These are the computational checks. They do not by themselves prove the model is adequate for interpretation.

Step 6: Extract the posterior

post <- get_posterior(fitted_model)

post is a tibble with one row per draw containing coef (named coefficient list), yhat (fitted values), noise_sd, r2, rmse, and smape.

Summarise coefficients:

library(dplyr); library(tidyr)

coef_summary <- post %>%
  select(coef) %>%
  unnest_wider(coef) %>%
  pivot_longer(everything(), names_to = "term") %>%
  group_by(term) %>%
  summarise(
    mean = mean(value), median = median(value), sd = sd(value),
    ci_low = quantile(value, 0.025), ci_high = quantile(value, 0.975),
    .groups = "drop"
  )
print(coef_summary, n = 30)

What to look for:

  • Media coefficients should be positive (boundaries enforce this).
  • Wide credible intervals mean the prior dominates, the data cannot identify the effect precisely.

Step 7: Assess model fit

fit_tbl <- fitted(fitted_model)
cat("Median R²:", median(r2(fitted_model)), "\n")
cat("Median RMSE:", median(rmse(fitted_model)), "\n")

For a well-specified MMM on weekly data, in-sample R² above 0.85 is typical.

Step 8: Response decomposition

tbl_map <- data.frame(
  actual.vars = c(
    "intercept", "events", "newsletters",
    "channel0_spend", "channel1_spend", "channel2_spend", "channel3_spend"
  ),
  group = c(rep("BASE", 3), rep("media", 4)),
  ref_points = c("none", rep("mean", 6))
)
decomp_result <- decomp(fitted_model, tbl_map = tbl_map)
head(decomp_result$DecompedData)

$DecompedData shows each term’s contribution (coefficient × design-matrix column) to the predicted KPI at each time point. The result is a native dsambayes_decomposition S3 object.

Step 9: MAP for rapid iteration

During development, use MAP for fast point estimates:

map_model <- model %>% fit_map(n_runs = 10)
get_posterior(map_model)

Use MCMC for final reporting; MAP for formula iteration. A MAP result is one point estimate, not posterior draws: it has no valid credible intervals or MCMC diagnostics. Review the restart diagnostics if objectives differ materially; use fit() when uncertainty, convergence, or decision risk matters.

Common pitfalls

Pitfall Symptom Fix
Forgetting to set boundaries Media coefficients go negative Add set_boundary(m_x > 0)
Too few iterations High Rhat, low ESS Increase iter and warmup
Missing controls High residual autocorrelation Add trend, seasonality, or holiday terms
Scaling confusion Coefficients look wrong model.scale: true is default; get_posterior() back-transforms automatically

Next steps

Your First Hierarchical Model

Goal

Build, fit, and interpret a multi-market hierarchical model with partial pooling and optional CRE (Mundlak) correction using the DSAMbayes R API.

This page is a hands-on tutorial. For the broader methodological questions behind prior-setting, diagnostics, and interpretation, use the workflow pages:

Prerequisites

Dataset

This walkthrough uses the tracked data/geo_panel/demo_data_geo_panel.csv , a panel dataset with weekly observations across multiple geos. Key columns:

  • Response: revenue, weekly synthetic KPI per geo.
  • Group: geo, geo identifier.
  • Media: channel0_signal, channel1_signal, channel2_signal, channel3_signal.
  • Controls: t_scaled, sin52_1, cos52_1, events, newsletters.
  • Date: date, weekly date index.
library(DSAMbayes)
panel_df <- read.csv("data/geo_panel/demo_data_geo_panel.csv")
table(panel_df$geo)  # Check group counts

Step 1: Construct the hierarchical model

The (term | group) syntax tells DSAMbayes to fit random effects. Terms inside the parentheses get group-specific deviations from the population mean:

model <- blm(
  revenue ~
    t_scaled + sin52_1 + cos52_1 + events + newsletters +
    channel0_signal + channel1_signal + channel2_signal + channel3_signal +
    (1 + channel0_signal + channel1_signal + channel2_signal + channel3_signal | geo),
  data = panel_df
)

This specifies:

  • Population (fixed) effects for all terms, the average effect across markets.
  • Random intercepts and slopes for media terms by market, each market can deviate from the population average.

Step 2: Set boundaries

model <- model %>%
  set_boundary(
    channel0_signal > 0, channel1_signal > 0,
    channel2_signal > 0, channel3_signal > 0
  )

Boundaries apply to the population-level coefficients.

Step 3: (Optional) Add CRE / Mundlak correction

If you suspect that group-level spending patterns are correlated with unobserved market characteristics (e.g. high-spend markets also have higher baseline demand), CRE can model this correlation when the conditional mean of the group effect is adequately represented by the included group means:

model <- model %>%
  set_cre(vars = c(
    "channel0_signal", "channel1_signal",
    "channel2_signal", "channel3_signal"
  ))

This adds one cre_mean_<channel> term for each media variable as a fixed effect. Conditional on the CRE specification and other included controls, the original media coefficient represents the within-group temporal association. CRE does not solve omitted time-varying confounding or establish a causal effect.

See CRE / Mundlak for when and why to use this.

Step 4: Fit with MCMC

fitted_model <- model %>%
  fit(cores = 4, iter = 2000, warmup = 1000, seed = 42)

Hierarchical models are slower than BLM, expect 10–30 minutes depending on group count and data size. First-time Stan compilation of the hierarchical template adds 2–3 minutes.

Step 5: Check diagnostics

chain_diagnostics(fitted_model)

Pay special attention to Rhat and ESS for sd_* parameters (group-level standard deviations), which are often harder to estimate than population coefficients.

Step 6: Extract the posterior

post <- get_posterior(fitted_model)

For hierarchical models, coefficient draws from get_posterior() return vectors (one value per group) rather than scalars. The population-level (fixed-effect) estimates are averaged across groups.

Step 7: Group-level results

Fitted values and decomposition are returned per group:

# Fitted values, one row per observation, grouped by geo
fit_tbl <- fitted(fitted_model)
head(fit_tbl)

# Decomposition, per-group predictor contributions
tbl_map <- data.frame(
  actual.vars = c(
    "intercept", "t_scaled", "sin52_1", "cos52_1", "events", "newsletters",
    "cre_mean_channel0_signal", "cre_mean_channel1_signal",
    "cre_mean_channel2_signal", "cre_mean_channel3_signal",
    "channel0_signal", "channel1_signal", "channel2_signal", "channel3_signal"
  ),
  group = c(rep("BASE", 10), rep("media", 4)),
  ref_points = c("none", rep("mean", 13))
)
decomp_result <- decomp(fitted_model, tbl_map = tbl_map)
# One native decomposition result per retained group:
head(decomp_result$decomp[[1]]$DecompedData)

Step 8: Budget optimisation (population level)

Budget optimisation uses population-level (fixed-effect) beta draws, not group-specific totals:

# See Budget Optimisation docs for full scenario specification
result <- optimise_budget(fitted_model, scenario = my_scenario)

Key differences from BLM

Aspect BLM Hierarchical
Data structure Single market Panel (multiple groups)
Coefficient draws Scalars Vectors (one per group)
Fit time 2–5 min 10–30 min
Decomposition Direct Native result per retained geo; probabilistic media transforms are rejected
Forest/prior-posterior plots Direct Group-averaged population estimates
Stan template bayes_lm_updater_revised.stan general_hierarchical.stan (templated per group count)

Common pitfalls

Pitfall Symptom Fix
Too few groups Weak partial pooling; group SDs poorly estimated Need 4+ groups for meaningful hierarchical structure
Too few obs per group High Rhat on sd_* parameters Increase iterations; simplify random-effect structure
CRE with too many vars More CRE variables than groups Reduce CRE variable set; see identification warnings
CRE mean has zero variance scale=TRUE aborts with constant column error Use model.type: re (without CRE) or model.scale: false

Next steps

Quickstart (YAML Runner)

Goal

Complete one reproducible DSAMbayes runner execution from validation to artefact inspection, then load the fitted model in R to explore the results interactively.

This page is operational by design. It teaches you how to run the package, not the full modelling methodology. After the quickstart succeeds, use Principled Bayesian Workflow before treating outputs as decision-ready.

Before you start

Complete the setup in Install and Setup. If you want to build a model interactively from R code instead of YAML, see Your First BLM Model.

1. Set up the environment

Complete the environment setup in Install and Setup (selecting the host library, creating the cache directory, and exporting XDG_CACHE_HOME) before continuing.

2. Validate the configuration (dry run)

Rscript scripts/dsambayes.R validate --config config/blm_timeseries.yaml

Expected: exits with code 0. No Stan compilation or sampling occurs.

3. Execute the full run

Rscript scripts/dsambayes.R run --config config/blm_timeseries.yaml

Expected: a timestamped run directory under results/ with staged outputs.

4. Locate and inspect the run directory

latest_run="$(ls -td results/* | head -n 1)"
echo "$latest_run"
find "$latest_run" -maxdepth 2 -type d | sort

Expected stage folders:

Folder Content
00_run_metadata/ Original/resolved/compiled configs, session info
10_pre_run/ VIF report, data dictionary, media spend plots
20_model_fit/ Fitted model object, fit plots
30_post_run/ Posterior summary, fitted/observed CSVs
40_diagnostics/ Diagnostics report, residual plots
50_model_selection/ LOO summary, Pareto-k diagnostics
60_scenario_analysis/ Posterior scenario summaries (when enabled)
70_forecast/ Reserved forecast stage (when enabled)
80_optimisation/ Budget allocation, response curves (when enabled)

5. Verify key artefacts

test -f "$latest_run/00_run_metadata/config.compiled.yaml" && echo "ok: config.compiled.yaml"
test -f "$latest_run/00_run_metadata/config.resolved.yaml" && echo "ok: config.resolved.yaml"
test -f "$latest_run/20_model_fit/model.rds" && echo "ok: model.rds"
test -f "$latest_run/30_post_run/posterior_summary.csv" && echo "ok: posterior_summary.csv"
test -f "$latest_run/40_diagnostics/diagnostics_report.csv" && echo "ok: diagnostics_report.csv"

6. Load the model in R

The fitted model is saved as an RDS object. Load it interactively to explore:

library(DSAMbayes)

model <- readRDS("results/<run_dir>/20_model_fit/model.rds")

# Posterior coefficient summary
post <- get_posterior(model)

# Fit quality
cat("Median R²:", median(r2(model)), "\n")

# Sampler diagnostics
chain_diagnostics(model)

# Fitted values
head(fitted(model))

7. Review diagnostics

Open 40_diagnostics/diagnostics_report.csv:

cat "$latest_run/40_diagnostics/diagnostics_summary.txt"

Quick interpretation:

  • pass: no immediate blocker
  • warn: review before sharing or acting
  • fail: do not treat the run as publishable or decision-ready

For the operational triage, see Interpret Diagnostics. For the methodological meaning of these gates, see:

8. Start from a tracked example config

Copy one of the two tracked examples and adapt it to your data:

  • config/blm_timeseries.yaml for single-series work
  • config/cre_geo_panel.yaml for geo-panel CRE work

Edit the copied YAML to point to your data and columns, then validate and run.

What the quickstart does not prove

A successful run means:

  • the package is installed correctly
  • the runner contract works on the example config
  • you have a complete staged result

It does not by itself prove:

  • the priors are appropriate
  • the fit is decision-ready
  • the decomposition is substantively meaningful
  • the optimisation output should be acted on

That is why the next stop should be the workflow pages.

If the quickstart fails

  • re-run validate before run
  • read the full error message
  • inspect 00_run_metadata/config.resolved.yaml
  • inspect 00_run_metadata/config.compiled.yaml
  • use Debug Run Failures

Next steps

FAQ

This page is for short, practical answers.

For the bigger methodological questions, start with:

Installation and setup

How long does the first Stan compilation take?

Usually 1 to 3 minutes. Subsequent runs typically reuse the cached binary. If compilation appears stuck, check the C++ toolchain in Install and Setup.

Do I need to set R_LIBS_USER every time?

Yes, unless you add it to your shell profile. Use dsambayes_set_r_library host to select a repo-local path that is isolated from both your system library and any container library.

Can I use renv instead of .Rlib?

Yes. The repo includes renv.lock. Use renv::restore() if you want exact dependency restoration.

Modelling

How many weeks of data do I need?

There is no hard minimum, but a useful rule of thumb is:

  • BLM: about 100+ weeks for a model with roughly 10 to 15 predictors
  • Hierarchical: about 80+ weeks per group, ideally with at least 4 groups

Shorter series can still be modelled, but the posterior will usually be much more prior-driven and less decision-ready.

Should I use identity or log response?

  • Identity when the KPI is naturally additive and variance is fairly stable
  • Log when the KPI is strictly positive and effect interpretation is more naturally multiplicative

If unsure, fit both and compare the adequacy and diagnostic picture, not just a single fit metric. See Response Scale Semantics.

How should I think about priors?

Start with Stage 2: Model and Priors.

The short answer is:

  • start with defaults
  • add sparse sign constraints only for structural assumptions
  • add sparse prior overrides only when you can defend them in business or modelling terms
  • do not use priors as a substitute for weak model design

Where do prior numbers come from, and why aren’t they standardised like some other MMM tools?

Short answer: in DSAMbayes, you write priors in the original units of your data, for example revenue, not in a standardised or z-score space. This is a deliberate design choice, not a mistake, and it differs from some other MMM tools, including Abacus, which ask you to specify priors directly in their internal standardised space.

What “standardised space” means

Before Stan samples, most Bayesian MMM tools, DSAMbayes included, centre and scale the response and predictors so the sampler works with unit-scale quantities. This is a computational convenience: it improves sampler efficiency and keeps step sizes well behaved across parameters with very different natural magnitudes (e.g. revenue in the hundreds of thousands versus a spend variable in the low thousands).

Where tools differ is which layer of that process they ask the user to think in.

  • Abacus-style: you write the prior directly on the standardised space. A config entry such as intercept: Normal(mu = 0, sigma = 2) is only interpretable because Abacus has already decided how the target is scaled; the 2 is “2 standard deviations of the standardised target”, not “2 units of revenue”.
  • DSAMbayes: you write the prior on the original, reporting-scale units of your data. intercept: Normal(mu = 120000, sigma = 15000) means “baseline revenue is around 120,000, plus or minus roughly 15,000”, in whatever currency or unit your target column uses. When model.scale: true (the default), DSAMbayes converts this raw-scale prior into the correct standardised parameterisation internally, fits, and then back-transforms the posterior draws to your original units automatically. You never see or touch the standardised numbers unless you set model.scale: false.

Both approaches produce equivalent underlying sampler behaviour. Neither is “unscaled” in the sense of skipping standardisation; DSAMbayes just does not require you to author priors in the post-scaling space.

Why DSAMbayes uses the reporting scale for the prior contract

  1. The numbers stay checkable by a non-modeller. A stakeholder, or a colleague reviewing your config, can look at mu = 120000, sigma = 15000 and immediately judge whether that is a sane assumption about baseline revenue. Nobody outside the modelling team can sanity-check Normal(0, 2) on a standardised scale without first recomputing what that implies in revenue terms.
  2. The prior’s meaning does not silently drift between runs. A standardised-space prior such as Normal(0, 2) means something different each time you refit, because the underlying standard deviation used to derive that scale changes with the data (different date range, different market, a new data pull). A reporting-scale prior such as Normal(120000, 15000) means the same thing, “baseline revenue is about 120k”, regardless of what the data’s standard deviation happens to be on any given run. This avoids a common failure mode where a prior that looked reasonable on one dataset becomes accidentally very informative, or accidentally very diffuse, on the next.
  3. It matches how priors are usually elicited in practice. Domain knowledge about a KPI baseline, a channel’s plausible ROI, or a competitor effect is almost always expressed in business units. Asking the analyst to translate that into a standardised space by hand is an extra, error-prone step that this design avoids.

This is consistent with standard Bayesian workflow guidance (see What Principled Means and Stage 2: Model and Priors): elicit priors where you have real, defensible intuition, and let the software handle re-parameterisation for computation.

What actually happens under the hood

DSAMbayes’s internal scaling of priors and boundaries, including the exact transformation ratios for slopes, the intercept, and hierarchical group-level standard deviations, is fully documented in Priors and Boundaries under “Scale semantics”. That page is the technical reference; this FAQ entry is only the “why”.

How to sanity-check a prior you have written

  • Reason about it directly in your data’s natural units; that is what the YAML priors: block and set_prior() both expect.
  • Use peek_prior(model) to see the full prior table on the reporting scale before fitting.
  • If you want to see the internal standardised-space numbers DSAMbayes actually passes to Stan, set model.scale: false on an otherwise identical model and compare, or inspect the scaling terms directly; this is rarely necessary for day-to-day modelling.
  • Do not manually convert your prior into a standardised space and enter that instead. That would be double-scaling and will produce the wrong prior.

See also Minimal-Prior Policy for the recommended default-first, sparse-override operating rule that this scaling behaviour is designed to support.

How many MCMC iterations do I need?

The defaults are a reasonable starting point. Then inspect the Stage 4 diagnostics:

  • Rhat <= 1.01 and healthy ESS usually mean the draw count is adequate
  • Rhat > 1.01 or weak ESS usually means you need to increase iterations and warmup
  • any divergences should be addressed before treating the fit as decision-ready

See Stage 4: Computation and Sampler.

How strict is the stationarity requirement for MMM?

DSAMbayes does not require the raw KPI to satisfy a textbook stationarity condition before fitting.

The important question is whether the remaining unexplained structure, after adding sensible controls and baseline terms, is weak enough that media effects are not standing in for missing baseline dynamics.

When should I set boundaries on media coefficients?

Use m_channel > 0 when non-negativity is a structural belief you would defend in writing. Do not apply blanket sign constraints just to make the output look tidier. See Stage 2: Model and Priors and Minimal-Prior Policy.

When should I use CRE (Mundlak)?

Use CRE when you want to separate within-group temporal effects from between-group cross-sectional structure in a hierarchical model. See CRE / Mundlak.

How should I handle CRE mean terms in decomposition / attribution?

Treat cre_mean_* terms as baseline or between-group structure, not as media attribution terms. They are there to absorb confounding structure, not to claim channel contribution.

What priors should I use on CRE mean terms?

Usually the defaults. Avoid manually tightening or positivity-constraining them unless you have a very strong reason, because that can undermine the whole point of CRE adjustment.

Can I add random slopes for CRE mean terms?

No. Those terms are constant within group, so random slopes on them are not separately identifiable from the group intercept.

What does scale = TRUE do?

It standardises the response and predictors before Stan fitting to improve sampler efficiency. Post-fit coefficient extraction is back-transformed automatically.

Runner and outputs

How long does a typical run take?

Roughly:

  • BLM MCMC: a few minutes
  • BLM MAP: seconds
  • Hierarchical MCMC: tens of minutes depending on size
  • Pooled MCMC: usually between BLM and hierarchical

First-time Stan compilation adds extra startup time.

What is the difference between validate and run?

  • validate checks config and data contracts without compiling or fitting Stan
  • run validates, fits, writes staged artefacts, and runs diagnostics

Always validate first after config changes.

Where do outputs go?

Under results/<timestamp>_<model_name>/ by default. See Output Artefacts.

How do I compare two model runs?

Use compare_runs() or compare the model-selection artefacts directly. See Compare Runs.

Diagnostics

Which diagnostics matter most?

Read them in order:

  1. Stage 4: Computation and Sampler
  2. Stage 5: Model Adequacy

That is more important than memorising one threshold in isolation.

What does “Pareto-k > 0.7” mean?

It means the LOO approximation is unreliable for that observation and the point is highly influential. Investigate the observation and be cautious about using LOO-based comparisons mechanically.

My diagnostics say warn. Should I worry?

Usually yes, but not always in the same way.

  • in exploratory work, a warning may be acceptable if understood
  • in shareable reporting, warnings should be disclosed and interpreted
  • repeated or severe warnings usually mean the model needs revision before decision use

Use Interpret Diagnostics for triage and the workflow pages for meaning.

Budget optimisation

How does the allocator work?

It searches feasible spend allocations within channel constraints and scores them against the fitted model. It is a decision layer built on the model, not an independent source of truth.

Can I use budget optimisation with MAP-fitted models?

Yes, but then the result is point-estimate-driven rather than uncertainty-rich. That is fine for rough iteration, not ideal for final decision support.