For Education Nonprofits & Funders
Your program collects scores, completion rates, attendance logs. But when a funder asks “did participants actually learn?” — can you decompose that claim into evidence?
The problem
A tutoring program reports “85% mastery rate.” A workforce training nonprofit reports “92% completion.” A digital learning app reports “100,000 sessions.”
But none of these numbers answer the question that matters: can participants do it on their own?
A student who got the right answer after a tutor walked them through it is not the same as a student who got the right answer alone. A completion certificate earned by clicking through modules is not evidence of skill. Reach is not impact.
Sound familiar?
Health nonprofit
Referrals up 35%. But engagement drops to 42%. For every 100 people contacted, 58 never become participants. The referral funnel leaks.
— Hegira Health, via Gaurav Mittal (Chronicle of Philanthropy)
Education program
“85% mastery rate.” But only 40% can perform without scaffolding. For every 100 students “mastered,” 60 can’t do it alone. The mastery funnel leaks.
— The problem YardStick solves
Digital health app
100,000 downloads. 2% retention. “A download is merely an invitation; true impact is measured by how many find the value to stay.”
— AHADI / TAHMEF, via Mittal
The measurement question
When a platform reports “mastery,” does that mean unaided performance on new problems — or correct answers after scaffolding? Most reporting doesn’t distinguish between the two.
— The question YardStick is designed to answer
What YardStick is designed to do
A correct answer after a hint carries less evidential weight than the same answer produced alone. The evidence trail records HOW each answer was produced — so it becomes possible to ask whether learning transfers when the help goes away.
Instead of a single pass/fail, YardStick evaluates four conditions independently: proficiency level, absence of misconceptions, performance on novel problems, and persistence over time. Each stream can pass or fail on its own — making it possible to see where a program's measurement breaks down.
Every mastery claim decomposes into the events that produced it. The measurement logic is not a black box — it is inspectable and verifiable. No vendor trust required. The goal is that programs can own their measurement, not rent it.
What you can verify today
We don’t have case studies yet — YardStick is new. What we can show you is the measurement itself working, live, in your browser:
Submit events with different scaffold contexts. Watch an unaided correct answer move the posterior 12x more than a helped one.
See how each false-mastery mode — fluent-but-flawed, memorizer, crammer — gets caught by a different gate stream.
Two learners, same correct-answer count. One was tutored. Without scaffold conditioning, they look identical. With it, the transfer gap is 71 points.
All four engines. Paste your own data into the calibration tool and get a verifiable certificate.
For funders
“When a student is reported as "mastered," can you show me the unaided evidence?”
If the answer is no, the mastery claim is unauditable. The program may be measuring tutor effort, not student learning.
“What percentage of your "mastered" students can perform on novel problems they haven't seen before?”
If the answer is unknown, the program can't distinguish memorization from understanding.
“How many students maintained mastery across multiple sessions over multiple weeks?”
If the answer is "we don't track that," the program can't distinguish cramming from durable learning.
No seat licenses. No vendor lock-in. Every claim decomposes into the evidence that produced it.