Config Schema

Purpose

This page documents the authored YAML contract used by:

  • scripts/dsambayes.R
  • DSAMbayes::run_from_yaml()
  • runme.R

The authored schema is schema_version: 2 only. Older formula-driven YAML files are intentionally rejected.

Processing order

The runner processes configs in this order:

  1. Parse YAML.
  2. Coerce YAML infinity tokens (.Inf, -.Inf).
  3. Apply v2 defaults.
  4. Resolve relative paths against the config file directory.
  5. Validate the authored v2 contract.
  6. Compile the authored config into the internal runner config.
  7. Apply managed holiday terms, then build the model and run.

Root sections

Key Required Purpose
schema_version yes Must be 2.
data yes Input data path, format, and date handling.
target yes Outcome column, KPI type, and response transform.
media yes Modeled media terms.
controls yes Non-media predictors, including manual trend/seasonality terms.
effects no Managed effects. In M1 this is holidays only.
model yes Model class and scaling options.
fixed_effects conditional Required exactly for model.type: fe.
hierarchy conditional Required for model.type: re and model.type: cre.
pooling conditional Required for model.type: pooled.
priors no Default priors plus grouped or explicit overrides.
boundaries no Grouped or explicit parameter boundaries.
fit no MCMC or optimise settings.
diagnostics no Diagnostics, model selection, and time-series selection settings.
allocation no Budget optimisation settings.
outputs no Output paths and artifact toggles.
forecast no Reserved forecast placeholder; currently only creates an empty stage directory when enabled.

Unknown keys fail validation.

Minimal valid config

schema_version: 2

data:
  path: ../data/timeseries/demo_data_synthetic.csv
  format: csv
  date_var: date

target:
  column: revenue
  type: revenue
  transform: identity

media:
  - channel0_signal
  - channel1_signal

controls:
  - t_scaled
  - sin52_1
  - cos52_1

model:
  type: blm

Key differences from the retired schema

  • model.formula is no longer authored directly.
  • schema_version: 1 configs are rejected.
  • Trend and seasonality stay user-authored as ordinary columns under controls.
  • Managed time effects are limited to holidays under effects.holidays.
  • re and cre models use hierarchy, not cre.enabled flags.
  • pooled models use pooling, not pooling.enabled.
  • fe models use fixed_effects.unit; hierarchy.group is not an FE alias.

Section reference

schema_version

Key Type Rules
schema_version integer Must be 2.

data

Key Type Rules
data.path string Required. File must exist. Relative paths resolve from the config directory.
data.format string csv, rds, or long.
data.date_var string Required in M1.
data.date_format string or null Optional parser format for date columns.
data.na_action string omit or error.
data.long_id_col string or null Required when data.format: long.
data.long_variable_col string or null Required when data.format: long.
data.long_value_col string or null Required when data.format: long.
data.dictionary_path string or null Optional metadata CSV.
data.dictionary mapping Optional inline metadata keyed by term name.

target

Key Type Rules
target.column string Required response column.
target.type string revenue or subscriptions.
target.transform string identity or log.
target.offset_column string or null Supported only for model.type: blm in M1.

media and controls

  • media is a required list of modeled media terms.
  • controls is a required list, but it may be empty ([]).
  • A term may not appear in both lists.
  • Manual trend and seasonality terms belong in controls.

All v2 names that become formula terms must be syntactic R names whose value is preserved by make.names(). This rule applies to target.column, target.offset_column, media, controls, fixed_effects.unit, hierarchy.group, and generated term prefixes. ASCII names such as sales_total and media.spend are the portable baseline. Non-ASCII names such as média are supported only when the active R locale preserves the exact name through make.names(). Rename columns containing spaces, operators, backticks, or backslashes before using them in a v2 config; examples such as sales value, paid-search, and sales`net are rejected during config validation. The formula sentinels ., ..., and ..1-style pronouns are also rejected because they do not evaluate as ordinary data columns.

Compiled formula order is:

  1. generated holiday terms
  2. controls
  3. media
  4. generated CRE mean terms
  5. optional offset
  6. hierarchical random-effects term

effects.holidays

Managed holidays are optional and are the only managed effect in M1.

effects:
  holidays:
    enabled: true
    path: ../data/holidays.csv
    label_col: holiday
    country: gb
    country_col: country
    week_start: monday
    prefix: holiday_
Key Type Rules
effects.holidays.enabled boolean Enables holiday feature generation.
effects.holidays.path string Required when enabled. CSV or RDS.
effects.holidays.date_col string or null Optional calendar date column override.
effects.holidays.label_col string Holiday label column.
effects.holidays.country string or null Optional single-country filter.
effects.holidays.country_col string Calendar column used with country.
effects.holidays.date_format string or null Optional parser format for non-ISO dates.
effects.holidays.week_start string monday through sunday.
effects.holidays.timezone string Timezone used in parsing/alignment. Must be a valid Olson timezone such as UTC.
effects.holidays.prefix string Prefix for generated holiday columns.
effects.holidays.window_before integer Non-negative.
effects.holidays.window_after integer Non-negative.
effects.holidays.aggregation_rule string count or any.
effects.holidays.overlap_policy string count_all or dedupe_label_date.
effects.holidays.overwrite_existing boolean Replaces existing columns only when true.

Notes:

  • The data date column must be aligned to the configured weekly anchor.
  • Country filtering materializes a filtered calendar artifact before the compiled config is written.

model

Key Type Rules
model.name string Defaults to the config filename stem.
model.type string blm, fe, re, cre, or pooled.
model.scale boolean Controls internal scaling before fit.
model.force_recompile boolean Forces Stan recompilation when true.

fixed_effects

fixed_effects is required exactly when model.type: fe. The approved shape has one unit key:

model:
  type: fe
  scale: true

fixed_effects:
  unit: geo

fixed_effects.unit must name one syntactic source column. It must differ from the target, date, media, and control columns. For long-format data it must equal data.long_id_col. FE configs cannot also contain hierarchy or pooling.

FE runner support includes validation, dry-run, and bounded non-dry MCMC fitting. Validation and dry-run construct and pre-flight the same unfitted fixed_effects model used by the direct API without Stan compilation or sampling. A materialised dry-run writes only 00_run_metadata/. A non-dry run uses the configured MCMC settings, the existing fit.fixed_effects() path, and the dedicated FE artefact writer.

The FE-safe resolved defaults are:

Field FE default or required value
data.na_action error
target.transform identity
fit.method mcmc
fit.mcmc.parameterization.positive_priors centered
diagnostics.model_selection.enabled false
diagnostics.time_series_selection.enabled false
diagnostics.identifiability.enabled false
diagnostics.enforce_publish_gate false
scenario_analysis.enabled false
allocation.enabled false
forecast.enabled false
outputs.layout staged

FE grouped priors support media_beta, control_beta, holiday_beta, and noise_sd. Explicit prior overrides may target an authored slope or noise_sd. Grouped boundaries support media_beta, control_beta, and holiday_beta; explicit boundary overrides may target authored slopes only. Intercept, CRE, pooling, random-effect, and residual-noise boundary requests are rejected.

FE v1 also rejects offsets, log targets, MAP, non-centred positive priors, internal media transformations, implicit row omission, scenario analysis, allocation, forecasting, model selection, time-series selection, generic identifiability output, publish-gate enforcement, flat output layout, level fitted or residual output, deployment, decomposition, and optimisation.

The resolved config exposes save_within_design_csv, save_contrast_residuals_csv, and save_contrast_ppc_csv for the dedicated FE artefact stage. Contrast-space fitted values are diagnostics, not level-scale fitted values or predictions. A completed FE run reports qualification_status: not_assessed and does not set a publishability result. No generic level-scale fitted, observed, residual, diagnostics, decomposition, scenario, optimisation, model-selection, time-series-selection, forecast, or deployment artefacts are written.

hierarchy

Required for model.type: re and model.type: cre.

Key Type Rules
hierarchy.group string Grouping column for panel models.
hierarchy.random_intercept boolean Include `(1
hierarchy.random_slopes list of strings Optional subset of authored media and controls.
hierarchy.cre_variables list of strings Required and non-empty for model.type: cre.
hierarchy.cre_prefix string Prefix for generated CRE mean terms. Default cre_mean_.

pooling

Required for model.type: pooled.

Key Type Rules
pooling.grouping_vars list of strings Required and non-empty.
pooling.map_path string Required. CSV or RDS.
pooling.map_format string csv or rds.
pooling.min_waves integer or null Optional positive integer.

priors

Key Type Rules
priors.use_defaults boolean Must remain true in M1.
priors.likelihood mapping Optional explicit alias for noise_sd.
priors.overrides list Explicit parameter-level overrides.

Grouped families are available when applicable:

  • intercept
  • media_beta
  • control_beta
  • holiday_beta
  • cre_beta
  • pooling_beta
  • random_effect_sd
  • noise_sd

Each grouped family accepts either the legacy DSAMbayes style:

family: normal    # or lognormal_ms where supported
mean: 0
sd: 0.5

or the more explicit alias:

distribution: Normal   # or HalfNormal / LogNormalMS where supported
mu: 0
sigma: 0.5

HalfNormal compiles to a zero-centered Normal prior plus an implied lower bound of 0 for unconstrained targeted parameter(s). Parameters that are already positive by construction, such as noise_sd and hierarchical sd_*[...], do not receive an extra boundary row.

The residual-noise prior also accepts this alias:

priors:
  likelihood:
    sigma:
      distribution: HalfNormal
      sigma: 2

boundaries

Boundary families mirror the grouped prior families and may also use explicit boundaries.overrides.

Each grouped or explicit boundary row uses:

lower: -Inf
upper: Inf

For FE, grouped boundaries support media_beta, control_beta, and holiday_beta. A boundary on noise_sd is not supported.

fit

Key Type Rules
fit.method string mcmc or optimise. Pooled runs require mcmc.
fit.seed numeric or null Optional scalar seed.
fit.optimise.* mapping Optimisation controls.
fit.mcmc.* mapping Stan sampling controls.
fit.mcmc.parameterization.positive_priors string centered or noncentered.

diagnostics

Retains the current runner surface for:

  • model_selection
  • time_series_selection
  • identifiability
  • publish-gate controls

Important M1 rule:

  • diagnostics.time_series_selection.enabled: true is not supported for pooled runs.
  • time-series selection is advisory only in the current release contract; it is not part of publish-gate enforcement.
  • lower-level runner paths with adstock/Hill media_transforms are not supported by time-series selection.
  • diagnostics.time_series_selection.gap_weeks is optional, defaults to 0, and inserts an embargo between the training window and the scored holdout window.
  • FE requires model selection, time-series selection, generic identifiability, and publish-gate enforcement to remain disabled.

scenario_analysis

Opt-in posterior scenario/reference evaluation after a successful MCMC fit:

scenario_analysis:
  enabled: true
  scenario_path: ../data/scenarios/planned.csv
  reference_path: ../data/scenarios/reference.csv
  scale: kpi
  log_response: mean
  interval: 0.9
  aggregate_by: []
  include_percent_lift: true
  save_draws: false
  • scenario_path and reference_path must identify CSV or RDS data frames. Relative paths resolve from the config file directory.
  • Both inputs must contain aligned rows and all predictors, offsets, date fields, and grouping fields required by the fitted model. They need not contain the response column.
  • fit.method must be mcmc. MAP output does not provide posterior contrasts.
  • scale is response or kpi. For log-response models, log_response chooses the lognormal conditional mean or the conditional median on the KPI scale.
  • aggregate_by names retained label columns. An empty list aggregates the full supplied path within each draw.
  • save_draws: false writes summaries and metadata only. Set it to true to add scenario_aggregate_draws.csv.
  • For fitted adstock/Hill media, the first supplied observation resets carry-over. If decision-horizon values depend on known earlier exposure, prepend the same observed history to both inputs. The runner reports every supplied row, so retain a field that distinguishes warm-up rows from decision rows if needed.

The output is a model-implied fitted-response contrast, not automatic causal attribution. See Counterfactual response.

allocation

Retains the current runner surface for budget optimisation, with channel targeting based on authored media terms.

  • allocation.n_candidates defaults to 2000 and must be a finite integer from 10 through .Machine$integer.max.
  • allocation.posterior.draws defaults to 500 and must be a finite integer from 1 through .Machine$integer.max.
  • Work scales approximately with candidates multiplied by retained posterior draws and, for efficient frontiers, by the number of feasible budget levels. See Budget Optimisation for measured review thresholds and the benchmark command.

outputs

outputs.root_dir and outputs.run_dir behave as before, but the metadata contract now includes:

  • config.original.yaml
  • config.resolved.yaml
  • config.compiled.yaml
  • outputs.save_model_rds controls the full fitted analysis artifact 20_model_fit/model.rds
  • outputs.save_deployment_model_rds controls the compact deployment artifact 20_model_fit/deployment_model.rds

When outputs.overwrite: true targets an existing run directory, the runner first validates the complete directory tree. It deletes only recognised flat or staged artifact paths. Unknown files, nested directories, and symbolic links cause an error before any existing artifact is deleted. Empty recognised stage directories may remain and are reused. Concurrent mutation of a run directory during overwrite is not supported.

Current first-slice limit:

  • outputs.save_deployment_model_rds: true is supported for model.type: blm, model.type: pooled with fit.method: mcmc, or hierarchical model.type: re / cre with fit.method: mcmc.
  • Pooled deployment artifacts score on authored terms and keep the normalized pooling map, but deployment-time newdata / data = ... does not need the pooling columns unless they are also ordinary formula terms.
  • Hierarchical deployment artifacts are seen-groups-only; explicit scoring/decomposition data must include the raw grouping columns, and decomposition also requires the response source column(s).
  • FE dry-runs write metadata only. Non-dry FE runs use bounded MCMC fitting and may write the dedicated posterior summary, sampler text, within-design, contrast-residual, and contrast posterior-predictive artefacts. They do not enter generic post-fit or decision-layer paths and remain unqualified.

forecast

Reserved placeholder only. In v1.3.3, enabling forecast can materialise 70_forecast/, but the runner does not emit forecast files or plots.

Examples in this repository

  • config/blm_timeseries.yaml, weekly time-series BLM example
  • config/fe_panel.yaml, weekly geo-panel FE validation, dry-run, and bounded MCMC example
  • config/re_geo_panel.yaml, weekly geo-panel RE example
  • config/cre_geo_panel.yaml, weekly geo-panel CRE example