Public-data calibration

Finding Percents

ASSISTments skill-builder dataset, 2012–13 public release · CML Rasch (log-domain ESF) · WLE person measures · calibrate-0.1.1 · seed 20260824

T1T2T3T4T5

All five gates passed · converged in 11 iterations · every number below recomputes from the released records

T1 |raw−model| = 0.011 · T2 bootstrap infit ∈ [0.935, 0.952] · T3 bank-median infit = 1.046 · T4 ledger reconciles · T5 engine−reference Δ = 0

certificate b3e1a70ae5e2a4a974aedd78f878d72ae8703ad2b784172d1779dc932f2dc9b9
212items calibrated
26,709 → 16,005raw → analyzed
1,455persons (non-extreme)
1.66separation · rel 0.735
+0.49targeting gap, logits
1.05 · 1.05infit · outfit medians

Fit figures are from the gated WO-HEFF-04 run; every item above the misfit threshold carries a one-line disposition in the report.

Hygiene ledger

Every removed row, named and reconciled. This table is gate T4.

StageResponses
Raw responses26,709
Scaffolding rows removed (original ≠ 1)−390
Duplicate student–item rows removed−1,591
Sparse items removed (< 50 responses)−3,732
Sparse persons removed (< 5 responses)−4,987
Degenerate items (p = 0 or 1)−4
Analyzed16,005 ✓

Misfit, dispositioned

No item appears here without an explanation. Gate T3: bank medians near 1, and every item above threshold carries a one-line disposition.

ItemN (informative)InfitOutfitpt-rDisposition
243101211.6213.790.72Explained: small N at the easy tail amplifies outlier residuals
1085310.0811.36−0.26Explained: N = 3 — the negative r is noise, not signal
39828136.3310.850.21Explained: small N at the easy tail amplifies outlier residuals
24307127.2510.350.34Explained: small N at the easy tail amplifies outlier residuals
593757505.468.130.22Explained: small N at the easy tail amplifies outlier residuals

47 of 212 items show outfit above 1.3 (22.2%) — a tail, not a pattern. All twelve dispositions across both skills are in the reports. The informative-N collapse traces to mastery-based stopping: items clear the 50-response floor on raw counts, then extreme-score exclusion leaves few informative responses at the tail.

Caught & fixed

Defects our gates found, in the open. This ledger is why the numbers above are trustable.

WO-HEFF-02 · targeting

Person measures received all 212 difficulties instead of only the items each person answered, deflating abilities on a 95%-sparse matrix. Caught by gate T1 (model-implied 2.7% vs observed 56.8%). Fixed; regression test pinned.

WO-HEFF-03 · fit statistics

Mean-square fit statistics violated the ≈1 invariant bank-wide. Caught by pre-release verification; gates T2, T3, and T5 added so the class cannot recur. Recalibrated under calibrate-0.1.1.

WO-HEFF-04 · close-out

Twelve extreme-misfit items individually dispositioned — informative N collapses at the tail once extreme-score students are excluded, making mastery-based stopping visible in the fit machinery. Gate T3 amended to bank medians plus per-item disposition; T1 spec pinned; certificate consistency enforced by the release script.

Recompute it yourself

pip install qlm-measure
qlm-measure verify finding_percents_report.json \
  --cert b3e1a70ae5e2a4a974aedd78f878d72ae8703ad2b784172d1779dc932f2dc9b9

Claims discipline: this page describes item properties of a historical public dataset only. No statements are made about student trajectories, instructional effectiveness, or current ASSISTments content. Data © ASSISTments 2012–13 public release, used under its research terms. CC BY 4.0 for our reports and parameters.