Registration Ledger

Pre-registered predictions alongside pilot results. Hypotheses committed before outcomes were known. The predictions include the engine losing its state-recovery comparison.

Frozen artifacts

spec_constants.yaml
8f8b63c4570653f5...48e2c
code tag
sim2-reg-v1
git hash
5b1472c615360f68...5baba
pipeline
v0.2-build1

Pilot disclosure

  1. BKT exceeded the engine on state-recovery AUC in all 36 cells, by approximately 0.01-0.03.
  2. Tempered-minus-untempered calibration error was negative in all 36 cells.
  3. False-verification rate was 1.0-2.0% in all cells.
  4. All four experimental factors moved their designated manipulation-check observables.

Pilot data excluded from all confirmatory analyses.

E1: False Verification Rate (auditability)

Hypothesis: FVR < 5% in every cell.

Decision: Pass iff point estimate < 0.05 in all 36 cells.

Pilot signal: 1.0-2.0% across all pilot cells.

E2: Calibration level

Hypothesis: Engine raw ECE < 0.20 in all medium- and rich-evidence cells.

Decision: Pass iff all 24 medium/rich cells are below 0.20. Sparse cells reported as directional boundary.

Pilot signal: ECE 0.10-0.15 in medium/rich; sparse is the stress test.

E4: State-recovery comparison

Hypothesis: Best-of-BKT >= engine AUC in a majority of cells (direction runs against the engine).

Decision: Report the paired cell-level AUC difference with 95% CI. No pass/fail gate.

Pilot signal: BKT exceeded engine by 0.01-0.03 in all 36 pilot cells.

E5: Tempering (calibration mechanism)

Hypothesis: Tempered minus untempered ECE < 0 in at least 33 of 36 cells, mean delta <= -0.03.

Decision: If H5 fails, tempering does not survive to engine v2 — explicit kill criterion.

Pilot signal: Negative in all 36 pilot cells.

Deviations

No deviations recorded. Any departure from this document will be logged here with date, reason, and assessed impact on each affected endpoint.