YardStick

The measurement platform whose instrumentation is open
and whose evidence trail is independently verifiable.

Researchers get a lab they can audit, not a vendor they must trust.

Start a Study -- FreeView Engine Validation
30/30
golden tests
match scipy / R
8/8
published studies
recovered
5
study types
verified end-to-end
170
tests passing
OneBastion suite
How It Works

Three phases. Full audit trail.

Every study follows the same rigorous pipeline -- from design to independently verifiable results.

1

Design

Power calculator, pre-registration, primary outcome freeze. Server-enforced -- no post-hoc outcome switching.

RCT Designer interface
RCT Designer with power calculator
2

Collect

Four outcome pipes: annotation, telemetry, qlm-measure SDK, and CSV upload. Pseudonymous by default.

Study enrollment page
Participant enrollment flow
3

Verify

Hash-chained records, standalone verifier, deterministic replay. Any third party can audit.

Engine validation dashboard
Engine validation with 8/8 checks
The Proof

8 Published Studies Recovered

Synthetic data calibrated to published parameters. Every engine recovers the known effect. All data is simulated -- this validates the statistical pipeline, not empirical claims.

Check PassedSimulated Data

Impasse-Driven Tutoring

VanLehn, K. (2011) -- Educational Psychologist, 46(4)
Published d = 0.79
Recovered d = 1.003
Check PassedSimulated Data

Process vs. Person Praise

Mueller, C. M. & Dweck, C. S. (1998) -- J. Personality & Social Psychology, 75(1)
Published: Process > Person
Recovered d = 2.20
Check PassedSimulated Data

Visible Learning Rankings

Hattie, J. (2009) -- Visible Learning. Routledge.
Published: FA > Clarity > Scaffolding
Recovered rho = 0.83
Check PassedSimulated Data

ICAP Framework

Chi, M. T. H. (2009) -- Topics in Cognitive Science, 1(1)
Published: Interactive > Active
Recovered: 27.4% mediated
Check PassedSimulated Data

High-Dosage Tutoring Meta-Analysis

Kraft, M. A. & Falken, A. (2021) -- EdWorkingPaper No. 20-335
Published d = 0.37
Recovered: N validated
Check PassedSimulated Data

Logistic Regression + AUC

Infrastructure Validation (2026) -- Internal validation
Expected AUC > 0.70
AUC validated
Check PassedSimulated Data

Chi-Square Independence

Infrastructure Validation (2026) -- Internal validation
Expected: reject null
Chi-sq validated
Check PassedSimulated Data

3-Arm Cluster Randomization

Infrastructure Validation (2026) -- Internal validation
Expected: balanced arms
Cluster integrity verified
Research-Grade Features

Built for journals, not dashboards.

Every feature exists because a methodologist asked for it.

Pre-Registration

Primary outcome frozen before data collection. Server-enforced -- no post-hoc outcome switching is possible once enrollment begins.

Bonferroni Correction

Secondary outcomes are automatically corrected for multiple comparisons. Exploratory analyses are labeled as such.

Participant-Level Nesting

N = participants, not sessions. Nesting is modeled. Cluster randomization handles classroom-level assignment correctly.

Hash-Chained Records

Every record carries a SHA-256 hash of its predecessor. Tampering breaks the chain. Any auditor can verify independently.

Power Refusal

Designs below 50% power are refused, not blessed. The system will not let you run an underpowered study and call it evidence.

APA Export

One-click publication-ready results with confidence intervals, effect sizes, and CONSORT flow diagrams. Ready for Methods sections.

Why It Matters

The Two Clocks

Speed and rigor are not tradeoffs. YardStick delivers both.

t_design

Question to pre-registered launch

A study that takes a semester elsewhere takes a week here. Power calculation, protocol design, pre-registration, and enrollment -- all in one session.

t_result

First enrollment to first result

Fast AND auditable -- speed a third party can verify. Hash-chained records mean velocity does not come at the cost of trustworthiness.

The Platform

See it in action.

Real product screenshots -- not mockups.

RCT Designer landing page
Design your study
Engine validation dashboard
Validate your engines
Simulated benchmarks page
Simulated benchmark results
Practice leaderboard
Track practice and progress

Start measuring what matters.

YardStick is free for research use.
No dark patterns. Pseudonymous by default.

Primary outcomes freeze before data collection -- by design.

Start a StudyView Engine Validation
YardStick by Quantum Learning Machines. All benchmark data is simulated.