Testing and Validation

Purpose

Define the canonical testing and validation workflow for the current DSAMbayes release (v1.3.5), from local pre-merge checks through release-quality gates.

Audience

  • Engineers running local checks before merge
  • Maintainers preparing release candidates
  • Reviewers validating release evidence

Validation layers

Layer Objective Primary command(s) Output proof
Lint Catch style and static issues early Rscript scripts/check.R --lint Exit code 0, no lint failures
Style Enforce formatting compliance on changed files Rscript scripts/check.R --style Exit code 0, no reformat-required files
Unit tests Catch behavioural regressions in package logic Rscript scripts/check.R --test Exit code 0, no test failures
Minimal smoke Keep a cheap Stan-backed safety check in the routine local loop Rscript scripts/check.R --smoke Exit code 0, fast unit tests plus minimal Stan smoke pass
Stan smoke Exercise the broader compiled Stan suite beyond the minimal smoke tier Rscript scripts/check.R --stan-smoke Exit code 0, Stan smoke tests enabled
Stan recovery evidence Re-run the full coefficient-recovery file under the nightly Stan gate Rscript scripts/check.R --stan-recovery Exit code 0, full-Stan DGP recovery tests pass
Stan release evidence Re-run high-budget pooled and hierarchical evidence checks for release candidates Rscript scripts/check.R --stan-release-evidence Exit code 0, targeted high-budget Stan evidence passes diagnostic thresholds
Docs sanity Catch local docs-link breakage and high-value contract drift Rscript scripts/check.R --docs Exit code 0, docs sanity checks pass
Package check Validate package-level install and check behaviour R -q -e 'rcmdcheck::rcmdcheck(...)' No ERROR; no unresolved WARNING
Runner validate Validate config and data contracts without fitting Rscript scripts/dsambayes.R validate ... Exit code 0, metadata artefacts
Runner run Validate end-to-end runner execution and artefacts Rscript scripts/dsambayes.R run ... Exit code 0, core run artefacts
Docs build Validate docs-site/Hugo buildability python3 docs-site/build_content.py && (cd docs-site && hugo --cleanDestinationDir) Exit code 0, successful site build

Environment setup

Run all commands from repository root:

# Navigate to your local DSAMbayes checkout
cd /path/to/DSAMbayes
source scripts/r-library-path.sh
dsambayes_set_r_library host
mkdir -p "$R_LIBS_USER" .cache
export XDG_CACHE_HOME="$PWD/.cache"

Expected outcome: checks run in a repo-scoped environment with reproducible library and cache paths.

Development dependency profile

The canonical development tools are declared in DESCRIPTION under Suggests and recorded in renv.lock: testthat, pkgload, lintr, styler, covr, and rcmdcheck. Restore this profile into the selected ABI-safe library without forcing optional modelling integrations:

R -q -e 'renv::restore(packages = c("testthat", "pkgload", "lintr", "styler", "covr", "rcmdcheck"), library = Sys.getenv("R_LIBS_USER"), prompt = FALSE)'

This targeted restore does not replace the release dependency-source check for private Git packages or a clean restore under the R version recorded in renv.lock.

Dependency source portability

Before a release candidate is signed off, verify all non-CRAN dependency sources in renv.lock and DESCRIPTION.

DSAMbayes owns the decomposition implementation. DSAMdecomp and teller are not package or lockfile dependencies. Release evidence for decomposition must record the DSAMbayes candidate commit and the native decomposition test results; do not record credentials in release evidence.

Local validation workflows

Developer fast path (pre-merge)

Use the explicit local ladder:

Rscript scripts/check.R --test
Rscript scripts/check.R --smoke
Rscript scripts/check.R --docs
Rscript scripts/check.R --release

Expected outcomes:

  • --test stays fast and runs the unit path only.
  • --smoke runs the unit path plus a minimal Stan-backed subset: one cheap BLM MCMC smoke, one pooled deployment-artifact smoke on a real pooled fit, and tiny run_from_yaml() runner smokes including pooled deployment_model.rds roundtrip coverage.
  • --docs runs the local docs/link sanity helper against README and the tracked docs surfaces.
  • --release runs lint, style, unit tests, minimal Stan smoke, docs sanity, and coverage as a broader local code gate. Coverage below the current 20% hard floor fails the lane.

Implementation note:

  • scripts/check.R --all remains a legacy convenience gate for lint, style, tests, and coverage only.
  • scripts/check.R --docs is the explicit docs/link drift check.
  • scripts/check.R --release is the clearer code-focused local release profile.
  • Neither profile replaces rcmdcheck, runner smoke checks, or docs build.

Stan-specific note:

  • use Rscript scripts/check.R --stan-smoke for the broader opt-in compiled Stan suite
  • use --stan-smoke-full for the fuller nightly variant
  • use --stan-recovery when you need a targeted rerun of the DGP recovery evidence without invoking the rest of the full Stan suite
  • use --stan-release-evidence for the high-budget release-candidate lane covering the warning-prone pooled and hierarchical paths

TSCV operational benchmark

Use the source-only benchmark to measure the current sequential, full-MCMC time-series cross-validation (TSCV) path before proposing checkpointing or parallel fold execution.

Inspect the fixed workload without loading the package, writing files, or running Stan:

Rscript scripts/benchmark_time_series_selection.R \
  --profile=quick --repetitions=1 --plan-only

Run the bounded BLM and correlated-random-effects (CRE) smoke profile:

Rscript scripts/benchmark_time_series_selection.R \
  --profile=quick --repetitions=1

Collect repeated one- and four-fold evidence for both model classes:

Rscript scripts/benchmark_time_series_selection.R \
  --profile=review --repetitions=3

The named profiles fix the dataset sizes, folds, chains, iterations and cores. Only the profile, repetition count from 1 to 10, deterministic base seed, and output directory are configurable. The driver first performs one unmeasured one-fold cache warm-up per model class, then runs every measured case in a fresh R process. It writes raw worker rows, logs, GNU time output, warmup_results.csv, and benchmark_results.csv below an ignored results/benchmark_tscv_* directory. GNU time writes stable, locale-neutral elapsed-time and peak-RSS markers; missing or malformed timing evidence fails the worker.

For review-profile comparisons, one- and four-fold cases of the same model class and repetition use identical synthetic data but distinct deterministic fit seeds. The output records both seeds, the Git commit and dirty state, the benchmark script hash, R/rstan/StanHeaders and platform details, the CPU model, the cache root, and whether the worker compiled or restored a cached model. Warm-up provenance is retained separately in warmup_results.csv.

Interpret the timing fields separately:

  • process_elapsed_sec and peak_rss_kb cover the whole fresh worker process;
  • runner_elapsed_sec covers the public runner call inside that process;
  • fold_fit_sec and fold_score_sec sum only the TSCV fold-level operations.

The benchmark disables unrelated runner diagnostics but keeps TSCV enabled. Its sampler budgets are intentionally too small for inference or convergence assessment. A successful worker still requires runner exit status zero, an overall TSCV status of ok, and every scheduled fold to succeed.

Treat the results as host-specific operational evidence. Cache state, host load, R and Stan versions, and filesystem behaviour affect them. They do not establish convergence, predictive superiority, statistical validity or causal validity. Use representative workload measurements and interruption evidence to decide whether TSCV optimisation is needed. If that work proceeds, specify and test checkpoint identity, deterministic seed recovery, cache isolation and atomic fold writes before considering bounded process-level parallelism.

The fixed review geometry uses the final 100-week training window for the one-fold case. The four-fold cases use 88-, 92-, 96- and 100-week expanding training windows with non-overlapping four-week holdouts. The 52-week minimum is a feasibility threshold, not the realised training size. Cross-class timing also reflects different workloads: review BLM has 104 input rows, whereas review CRE has 624 rows and hierarchical structure. Do not interpret their difference as an intrinsic model-class multiplier.

Release-candidate full path

Run mandatory gates in this exact order:

Rscript scripts/check.R --lint
Rscript scripts/check.R --style
Rscript scripts/check.R --smoke
Rscript scripts/check.R --stan-release-evidence
_R_CHECK_FORCE_SUGGESTS_=false \
  R -q -e 'rcmdcheck::rcmdcheck(args = c("--no-manual"), error_on = "warning")'
Rscript scripts/dsambayes.R validate --config config/blm_timeseries.yaml --run-dir results/quality_gate_validate
Rscript scripts/dsambayes.R run --config config/blm_timeseries.yaml --run-dir results/quality_gate_run
python3 docs-site/build_content.py
(cd docs-site && hugo --cleanDestinationDir)

Expected outcome: all gates complete with exit code 0, with no unresolved release blockers.

Local package-check note:

  • Canonical local rcmdcheck runs set _R_CHECK_FORCE_SUGGESTS_=false. Native decomposition needs no external decomposition or telemetry package.

Runner smoke-test expectations

Minimum release smoke expectations:

  1. scripts/check.R --smoke succeeds, proving at least one cheap Stan MCMC path, pooled deployment-artifact coverage on a real pooled fit, and tiny runner fit paths.
  2. validate command succeeds and writes metadata artefacts.
  3. run command succeeds and writes model, fitted/observed output, and diagnostics artefacts.
  4. Required runner artefact paths exist under results/quality_gate_validate/ and results/quality_gate_run/.

For matrix and exact artefact paths, use Runner Smoke Tests.

Evidence capture requirements

Before sign-off, capture:

  1. Full command logs and exit codes for all mandatory gates.
  2. Runner smoke artefacts from validate and run directories.
  3. Candidate commit hash and top changelog section.

Use Release Evidence Pack as the authoritative bundle contract.

Failure handling

  1. Any gate failure is a release blocker until resolved.
  2. Re-run the full failed gate after remediation.
  3. If rcmdcheck emits NOTE, record reviewer rationale explicitly.
  4. If runner artefacts are missing, inspect resolved config and outputs.* flags.