Counterfactual Response

Purpose

counterfactual_response() evaluates one fitted MCMC model under two aligned predictor paths. It returns posterior draws of the fitted response location under the proposed scenario, the reference response, and their difference.

The result is a model-implied scenario contrast. It is not, by itself, a causal effect or an attribution estimate. Causal interpretation requires a credible identification argument for the fitted model and the intervention being represented.

Usage skeleton

The following code assumes that fitted_model is a fitted MCMC object and future_data contains every required model variable and modelling wave.

scenario <- future_data
scenario$paid_search <- scenario$paid_search * 1.10

reference <- future_data

contrast <- counterfactual_response(
  fitted_model,
  scenario = scenario,
  reference = reference,
  scale = "kpi",
  interval = 0.9
)

contrast$summary
contrast$draws

Both data frames must resolve to the same retained rows. Dates, hierarchical group keys and row order must match. Predictor values may differ. The reported difference is scenario - reference. For balanced panels, supply the common ordered modelling waves used by the fitted model.

Returned quantities

The object has three components:

Component Content
draws One row per selected posterior sample. The scenario, reference, and difference columns contain response vectors.
summary One row per retained observation with posterior means, medians, the requested central interval and .input_row provenance.
metadata Estimand, output scale, transformation method, draw count, carry-over initialisation and interpretation limits.

These are posterior draws of the fitted response location. They do not include new observation noise. On KPI scale, a log-response model can return a conditional median or a lognormal conditional mean. Use the draws to describe uncertainty in the fitted response surface, not the full range of future observations.

Use draw_ids to select existing .sample identifiers when the complete draw-by-row result would be too large:

contrast <- counterfactual_response(
  fitted_model,
  scenario,
  reference,
  draw_ids = seq(1, 2000, by = 4)
)

draw_ids reduces the returned draw-by-row object. The current implementation still materialises the fitted posterior before selecting those identifiers, so it does not remove the peak memory cost of posterior extraction.

Aggregate a decision horizon

Use aggregate_counterfactual_response() to sum responses within each posterior draw before calculating intervals:

horizon <- aggregate_counterfactual_response(contrast)

horizon$summary
horizon$draws

The result includes draw-wise scenario totals, reference totals, differences and percentage lift. Its summary includes posterior means, medians, central intervals and probability_difference_positive. That probability is the proportion of retained posterior draws where the aggregated scenario - reference difference is strictly above zero. It is not the probability that an intervention has a causal effect.

Group by a retained panel or date label when separate aggregates are needed:

by_market <- aggregate_counterfactual_response(contrast, by = "market")
by_wave <- aggregate_counterfactual_response(contrast, by = "date")

Aggregation is draw-wise. DSAMbayes first sums each draw over the relevant rows and then calculates posterior intervals. It does not sum row-level interval bounds.

Percentage lift is defined draw by draw as:

100 * difference_total / reference_total

It requires totals on the KPI or identity-response scale and a strictly positive reference total in every retained draw and group. For log-response totals, first call counterfactual_response(..., scale = "kpi"). Set include_percent_lift = FALSE when only absolute totals and differences are required.

YAML runner integration

Set scenario_analysis.enabled: true to evaluate two file-backed paths after a successful MCMC fit. The runner writes row summaries, aggregate summaries and metadata under 60_scenario_analysis/. Draw-level aggregate output is written only when scenario_analysis.save_draws: true.

scenario_analysis:
  enabled: true
  scenario_path: ../data/scenarios/planned.csv
  reference_path: ../data/scenarios/reference.csv
  scale: kpi
  log_response: mean
  interval: 0.9
  aggregate_by: []
  include_percent_lift: true
  save_draws: false

The scenario and reference files must follow the same alignment and input contracts described above. A failed requested scenario analysis does not erase the successful model fit. The returned runner result instead records outcome = "completed_with_scenario_analysis_fail" and the failure reason. See Config Schema and Output Artefacts.

Response and KPI scales

scale = "response" returns the model response scale. For log(y) models, this is the log scale.

scale = "kpi" converts both absolute paths before calculating their difference. For a log-response model:

  • log_response = "median" uses exp(eta);
  • log_response = "mean" uses the draw-specific lognormal adjustment exp(eta + noise_sd^2 / 2).

The second form estimates a conditional expected KPI level under the fitted Gaussian log-response model. It does not correct model misspecification or confer a causal interpretation.

Media transforms and carry-over

For a model fitted with probabilistic adstock and Hill saturation, the engine uses each posterior draw’s fitted decay, half-saturation and media coefficient. It uses the fixed Hill shape and scaling metadata stored with the fit. The full Stan posterior must therefore remain available. A materialised posterior table or compact deployment artefact does not contain enough information and is rejected.

Carry-over starts at the first supplied row. DSAMbayes does not invent media history before the scenario window. Hierarchical media transforms reset at the fitted panel boundary. The current pooled Stan model uses one media panel, so its carry-over continues across the complete supplied row sequence. Arrange pooled inputs to match the fitted row contract and review this limitation before using pooled transformed-media contrasts.

If the scenario continues from a known observed media history, prepend that common history to both scenario and reference. Keep the historical values identical in both paths. This reconstructs carry-over at the decision-horizon boundary. After evaluation, use .input_row or the date and panel keys to keep only the decision-horizon rows. Omitting known pre-horizon media resets carry-over and answers a different scenario question.

Changing spend early in the supplied window can affect later rows through carry-over. Scenario and reference paths should therefore cover the complete decision horizon rather than isolated rows.

Hierarchical paths are evaluated in the model’s group-major order. The summary$.input_row column maps every processed row back to its supplied data row. Draw-vector positions follow summary$.row; do not bind them back to the input data by position alone.

Supported and rejected fits

The first implementation supports fitted MCMC BLM, hierarchical and pooled models. It rejects:

  • unfitted models;
  • MAP fits, because one optimum is not a posterior sample;
  • deployment artefacts, because they contain point estimates;
  • transformed-media objects that no longer retain Stan alpha and k draws;
  • unseen hierarchical levels;
  • duplicate date/panel row keys;
  • transformed paths that are not strictly ordered by modelling wave;
  • transformed hierarchical paths whose panels do not contain identical wave sets;
  • scenario and reference paths with different retained dates, panels or row counts.

optimise_budget() remains a separate scenario-authored response-surface workflow. It does not yet call this engine and does not automatically inherit fitted adstock or Hill parameters.