Technical Documentation

Quality Framework

Structured under the field’s five quality criteria. Validity follows Pellegrino’s three-part structure. Each heading has either evidence or a stated gap.

1. Content & Context

Domain coverage

Ontology-based construct graph. 30-construct synthetic DAG with fragility assignment (durable/fragile/stubborn). Clinical packs: 19 CLIN rules, 9 LabPath rules.

Evidence

Scaffold taxonomy

Five-level hierarchy: unaided (w=0.95) → open question → hint → worked example → direct instruction (w=0.08). Weight ordering verified in patent golden numerics.

Evidence

Misconception model

Distractor-targeted generation with misconception-bound traps (CLIN-01). Cascade DAG for misconception propagation. Resolution durability tracking.

Evidence

Content authoring

Wrightyard validator harness: 28 executable rules, severity staging, fail-closed on crash. Case Source Compiler with trace manifests.

Evidence

2. Scientific Soundness

Validity · Reliability · Comparability

Cognitive validity (Pellegrino)

Construct graph grounds measurement in domain structure. Depth levels (procedural → conceptual → transfer → integration) map to cognitive complexity.

Evidence

Instructional validity (Pellegrino)

Scaffold-conditioned weight links measurement to instruction. tierBlockedBy drives next-action recommendations. Policy module maps depth to allowed scaffolds.

Evidence

Inferential validity (Pellegrino)

Tempered evidence update verified against patent golden numerics (0.5032, 0.8087). CML Rasch calibration with self-test harness (4 designs, all converge, correlation >0.97). 36-cell SIM-001 audit with pre-registered predictions. GAP: No field trial data yet. Inferential validity against human outcomes not established.

Partial

Reliability

Deterministic replay: same seed → same hash. Certificate tamper detection via SHA-256 chain. Self-test harness proves recovery at >95% coverage.

Evidence

Comparability

CRN (common random numbers) in SIM-001: base uniforms drawn once, mapped through cell-specific quantile functions. Seeded randomness ensures cross-cell comparability.

Evidence

3. Usefulness

Decision-tier tagging

Four tiers (instructional → reporting → placement → credential) with ascending evidence requirements. tierBlockedBy tells consumers what stream is short.

Evidence

False-mastery detection

4-stream mastery gate catches fluent-but-flawed (misconceptions), memorizers (no transfer), crammers (no persistence). Each mode fails exactly one stream.

Evidence

Scaffold-conditioned measurement

Distinguishes helped from unaided performance. Side-by-side comparison demo: same correct rate, 71-point transfer gap. Help is context, not penalty.

Evidence

Audit harness

36-cell factorial with pre-registered predictions including negative results. Engine loses on state recovery (pre-registered). Wins on calibration and auditability.

Evidence

4. Usability

Learner view

Evidence chain rendered in plain language. Scaffold context as fact, never penalty. No numeric weights or thresholds. Family variant with zero internal vocabulary.

Evidence

Teacher view

Full evidence ledger with context tags. P(mastery) with depth distribution. Certification card with 4-stream gate checks. Countersign flow.

Evidence

Researcher interface

Interactive demos (sandbox, gate, compare). Workshop mode. Live calibration API. 8 published studies recovered.

Evidence

Non-use guardrails

permittedUses/prohibitedUses on every record. Default prohibits educator evaluation. Consent-scope gated exports. Withdrawal cascade.

Evidence

Accessibility

Keyboard parity declared in Wrightyard rules (CLIN-13) but not yet audited on the YardStick product UI itself. WCAG audit not performed.

Gap

5. Scalability

Architecture

Standalone repo, independent Vercel + Render deployment. 353 engine tests, 54 work-order tests, all passing independently.

Evidence

Calibration at scale

CML Rasch + PCM with Newton-Raphson convergence. N=500 factorial (36 cells) running. Privacy export with sufficient-statistic mode.

Evidence

API

23 endpoints. Health, events, state, evidence, certifications, calibration, audit, wrightyard. All verified on production.

Evidence

Multi-tenant

Current deployment is single-tenant. Multi-district isolation not implemented. Consent-scope enforcement is per-record but not per-tenant.

Gap

Patent Basis

U.S. Patent Application No. 19/709,832 (filed June 16, 2026). Controlled Multi-Agent Artificial Intelligence Inquiry System with Coauthorship Integrity Verification, Escalating Reasoning Pressure Protocol, Adaptive Misconception-Targeted Distractor Generation, and AI-Resistant Behavioral Evidence Capture. Inventor: Kumar Sumbhav Srivastava.