For Education Researchers
When a platform claims a student “mastered” a skill, can you decompose that claim into evidence? Can you tell whether mastery was demonstrated independently or with help? Can a reviewer reproduce the measurement?
The problem
You evaluate an intervention — a tutoring program, a teacher training model, a policy change — by measuring student outcomes. But the measurement itself is opaque.
Standardized test scores are coarse. Adaptive platform claims about “mastery” are unauditable. When you publish, reviewers ask “how do you know the student actually learned this?” and the honest answer is: the vendor said so.
The treatment effect is only as good as the outcome measure. If the measure can’t be verified, neither can the finding.
What YardStick gives you
When an adaptive platform claims a student mastered a skill, you can decompose that claim into the evidence trail. Was it unaided performance or help-assisted? How many sessions? Did it hold over time? This matters enormously for intervention studies — the treatment effect is only as good as the outcome measure.
Current measures can't distinguish "the student got the right answer alone" from "the student got the right answer after the tutor walked them through it." YardStick can. That's the difference between measuring actual learning transfer and measuring compliance.
No post-hoc outcome switching. The primary outcome is frozen before enrollment. This is what journals increasingly demand, and it's enforced by the platform — not by the researcher's self-discipline.
When you build or use an assessment, the calibration produces a signed certificate anyone can reproduce. The self-test harness proves the estimator recovers known parameters. No "trust our software" — check it yourself.
A study that takes a semester to design and instrument takes a week on YardStick. Power calculation, protocol design, pre-registration, enrollment — one session. The speed is auditable: hash-chained records mean velocity does not come at the cost of trustworthiness.
Example
You’re running an RCT on a tutoring intervention. The treatment group gets 3x/week tutoring sessions. The control group gets business-as-usual instruction.
Without YardStick: You measure outcomes with a pre/post standardized test. The test tells you the treatment group scored 0.3 SD higher. But you can’t tell whether that gain came from genuine conceptual understanding or from tutors coaching students through similar problems. The measure is agnostic to howthe answer was produced.
With YardStick: Every practice response carries its scaffold context. An unaided correct answer on a novel problem carries full evidential weight. A correct answer after a hint carries less. Mastery is certified only when a student passes on unaided novel probes, across multiple sessions, over 7+ days. You can report not just that the treatment group improved, but how — and whether the improvement transfers when the tutor isn’t present.
Try it now
Walk through all 7 capabilities with synthetic data. No sign-up required.
Open demo →
Guided hands-on walkthrough. Paste your own data, get live results from the API.
Start workshop →
The Open Measurement Layer — why open education AI needs one more layer.
Read manifesto →
Architecture, validation, and the case for separating measurement from instruction.
Read paper →
No seat licenses. No dark patterns. Pseudonymous by default. Primary outcomes freeze before data collection — by design.