Person measures received all 212 difficulties instead of only the items each person answered, deflating abilities on a 95%-sparse matrix. Caught by gate T1 (model-implied 2.7% vs observed 56.8%). Fixed; regression test pinned.
Public-data calibration
Finding Percents
ASSISTments skill-builder dataset, 2012–13 public release · CML Rasch (log-domain ESF) · WLE person measures · calibrate-0.1.1 · seed 20260824
All five gates passed · converged in 11 iterations · every number below recomputes from the released records
T1 |raw−model| = 0.011 · T2 bootstrap infit ∈ [0.935, 0.952] · T3 bank-median infit = 1.046 · T4 ledger reconciles · T5 engine−reference Δ = 0
Fit figures are from the gated WO-HEFF-04 run; every item above the misfit threshold carries a one-line disposition in the report.
Hygiene ledger
Every removed row, named and reconciled. This table is gate T4.
| Stage | Responses |
|---|---|
| Raw responses | 26,709 |
| Scaffolding rows removed (original ≠ 1) | −390 |
| Duplicate student–item rows removed | −1,591 |
| Sparse items removed (< 50 responses) | −3,732 |
| Sparse persons removed (< 5 responses) | −4,987 |
| Degenerate items (p = 0 or 1) | −4 |
| Analyzed | 16,005 ✓ |
Misfit, dispositioned
No item appears here without an explanation. Gate T3: bank medians near 1, and every item above threshold carries a one-line disposition.
| Item | N (informative) | Infit | Outfit | pt-r | Disposition |
|---|---|---|---|---|---|
| 24310 | 12 | 11.62 | 13.79 | 0.72 | Explained: small N at the easy tail amplifies outlier residuals |
| 1085 | 3 | 10.08 | 11.36 | −0.26 | Explained: N = 3 — the negative r is noise, not signal |
| 39828 | 13 | 6.33 | 10.85 | 0.21 | Explained: small N at the easy tail amplifies outlier residuals |
| 24307 | 12 | 7.25 | 10.35 | 0.34 | Explained: small N at the easy tail amplifies outlier residuals |
| 593757 | 50 | 5.46 | 8.13 | 0.22 | Explained: small N at the easy tail amplifies outlier residuals |
47 of 212 items show outfit above 1.3 (22.2%) — a tail, not a pattern. All twelve dispositions across both skills are in the reports. The informative-N collapse traces to mastery-based stopping: items clear the 50-response floor on raw counts, then extreme-score exclusion leaves few informative responses at the tail.
Caught & fixed
Defects our gates found, in the open. This ledger is why the numbers above are trustable.
Mean-square fit statistics violated the ≈1 invariant bank-wide. Caught by pre-release verification; gates T2, T3, and T5 added so the class cannot recur. Recalibrated under calibrate-0.1.1.
Twelve extreme-misfit items individually dispositioned — informative N collapses at the tail once extreme-score students are excluded, making mastery-based stopping visible in the fit machinery. Gate T3 amended to bank medians plus per-item disposition; T1 spec pinned; certificate consistency enforced by the release script.
Recompute it yourself
pip install qlm-measure qlm-measure verify finding_percents_report.json \ --cert b3e1a70ae5e2a4a974aedd78f878d72ae8703ad2b784172d1779dc932f2dc9b9
Claims discipline: this page describes item properties of a historical public dataset only. No statements are made about student trajectories, instructional effectiveness, or current ASSISTments content. Data © ASSISTments 2012–13 public release, used under its research terms. CC BY 4.0 for our reports and parameters.