CascadeBench

A benchmark is valid only if tomorrow never leaks into yesterday.

CascadeBench freezes a decision cutoff, separates model inputs from post-event outcomes, and blocks scoring when temporal or source partitions fail.

Reference cases12
Historically scored0
External validations0
Real-user impact studies0
empirical validation not yet established

The twelve launch cases verify execution, invariants, evidence governance, packaging, and tamper detection. They do not support a claim of predictive accuracy.

Mean absolute error

Magnitude agreement between central scenario impact and a comparable separated outcome proxy.

Spearman rank

Whether node ordering agrees, without pretending that rank correlation establishes causal validity.

Interval coverage

Share of separated outcome observations contained by the declared lower–upper envelope.

Coverage calibration

Absolute gap between declared full-envelope coverage and empirical coverage, reported with mean interval width.

Direction accuracy

Whether impact exceeds a frozen materiality threshold in both prediction and outcome.

Regret vs. zero baseline

Excess absolute error over an explicit no-impact baseline; negative values mean the model improves on that baseline.

Leakage audit

Input availability, observation time, source role, and outcome partition are checked before scoring.

Scenario-only fallback

Used when outcomes are missing, incomparable, too few, or the case is explicitly synthetic.

Minimum acceptable historical replay

  1. Freeze a decision cutoff before inspecting the evaluation outcomes.
  2. Preserve exact input artifacts and prove they were available by the cutoff.
  3. Acquire outcomes through a distinct source role after the event.
  4. Predeclare comparable nodes, proxy definition, horizon, threshold, exclusions, and missing-data policy.
  5. Publish the full RiskPack and retain failed or blocked cases in the denominator.

Open the runnable replay, review, user-study, adoption, and impact protocols ↗