Current status: Methodology version 1.0. Last reviewed August 16, 2026. No model or product scores are published on this page. Material changes will be recorded and affected evaluations will be rerun.
TrueMadeAI evaluates K-12 AI in five separate layers: product controls, district deployability, operational performance, pedagogical behavior, and actual learning outcomes. Every result must name the exact product, model, date, account, configuration, and test case. We do not turn these different questions into one universal winner, and we do not call a polished tutoring transcript proof that a student learned.
This page is the public protocol for future TrueMadeAI and Tenet scorecards. It contains no model rankings or test results. It explains what we will test, what evidence a score can support, and where a district must still conduct its own review.
The method draws from NIST’s expectation that AI evaluation be documented, repeatable, connected to deployment context, and explicit about uncertainty. It also incorporates the distinction in K-12EduBench between answer knowledge, problem solving, and educational-goal cognition; the educator-rating approach used in LearnLM research; and the learning-outcome warning demonstrated in a large high school mathematics field experiment.
The five evidence layers stay separate
An AI product can be fast and inexpensive but pedagogically weak. A model can produce excellent hints but be unavailable through a district-manageable account. A guarded tutor can preserve student effort better than a general chatbot even when both use the same underlying model. One combined score would hide these distinctions.
| Evidence layer | Question answered | Evidence used | What it cannot prove |
|---|---|---|---|
| Product and account controls | Can the district configure and govern the product for the intended users? | Official terms and documentation, administrator settings, account tests, and support confirmation | That every model response will be correct or instructionally sound |
| District deployability | Can the district operate the tested configuration within its identity, device, privacy, accessibility, support, and procurement constraints? | Configuration record, deployment test, data-flow review, accessibility checks, and district requirements | That another plan, region, device, or contract has the same properties |
| Operational performance | How quickly, reliably, and economically does the configuration complete the defined task? | Timed runs, errors, retries, token counts where exposed, dated prices, throughput, and reviewer correction burden | That speed or low cost produces learning |
| Pedagogical behavior | Does the observed interaction behave like useful instruction for the stated goal? | Blinded transcripts scored by educators against a predeclared rubric | That a real student retained or transferred knowledge |
| Actual learning outcomes | Did an intervention change unassisted student learning? | An appropriately approved study with baseline, comparison, unassisted outcomes, and retention or transfer measures | That the result generalizes beyond the studied population, implementation, subject, and time period |
We publish domain results side by side. We do not average a privacy-documentation finding, a latency measurement, and an educator judgment into a single number.
Districts evaluating a purchase should pair scorecard evidence with the K-12 AI governance software buyer’s guide, the AI tool vetting and approval template, the K-12 AI Assistance Ladder, and their own legal, security, accessibility, instructional, and procurement review. The first application of this structure is the dated ChatGPT for Teachers vs Claude for Teachers vs Gemini for Education comparison.
The model is not the product
A model is the inference system that produces an output. A product can surround that model with system instructions, retrieval, search, tools, memory, filters, account eligibility, administrator controls, data retention, model routing, and user-interface choices. Those layers can materially change both the educational experience and the data path.
We therefore use two distinct tracks:
- Product UI track. We test the experience a student or educator would actually use under a named plan and account configuration. The result applies to that product experience, even when the provider does not expose an exact model snapshot.
- API or local-model track. We test a named endpoint or downloadable model with a frozen system prompt and inference configuration. The result offers better model-level reproducibility, but it does not inherit the controls or experience of a provider’s finished product.
We never transfer a result between tracks. A score from an API endpoint is not presented as a score for a consumer chat product, and a product-UI result is not presented as the inherent capability of a single underlying model.
What we record before the first run
Each evaluation receives a public configuration record. If a field is hidden by the provider, we label it not disclosed rather than guessing.
The record includes:
- evaluation date, start time, end time, and time zone;
- provider, product name, plan, account type, and intended user role;
- product URL and interface version when visible, or API endpoint and SDK version;
- the provider-reported model name, exact model snapshot or artifact hash when available, and any visible automatic-routing notice;
- system and developer instructions, user prompts, fixed follow-up turns, attachment names, and file hashes;
- whether the run used a clean new conversation, temporary chat, memory, history, personalization, search, retrieval, browsing, tools, or connected applications;
- API parameters such as temperature, top-p, seed, maximum output, and tool-choice settings when exposed;
- language, locale, device, browser, operating system, and relevant accessibility settings;
- for a downloadable model, weights hash, quantization, inference runtime, prompt template, accelerator, memory, operating system, and network state;
- administrator controls, age or account eligibility, retention and training-use settings, and safety configuration relevant to the test;
- list price and currency on the test date, plus input and output units when the interface exposes them;
- run identifier, raw transcript hash, screenshots or logs needed to reproduce an observed interface fact, and any operator deviation.
This configuration record prevents an ambiguous statement such as “Model X is better for schools.” The defensible statement is narrower: a dated configuration performed in a specified way on specified cases under a specified protocol.
For privacy and procurement review, the configuration record links to a separate K-12 AI data-boundary analysis. Provider statements about retention, training, contracts, subprocessors, or age eligibility remain attributed provider claims until independently observable or contractually verified.
Product and account control review
This review describes the product a district can actually procure and administer. It does not award a product the benefit of a control available only on a different plan or account type. The evidence table covers:
- student and educator eligibility, age limits, parental-consent responsibilities, and education-specific account availability;
- administrator roles, single sign-on, provisioning, deprovisioning, roster scope, and least-privilege options;
- training-use defaults and controls, retention, deletion, export, subprocessors, region, and available contract terms;
- product-level search, memory, file, image, voice, connector, sharing, and public-link controls;
- safety and academic-integrity settings, teacher visibility, exception paths, and incident reporting;
- accessibility documentation plus direct keyboard, screen-reader, contrast, caption, and error-recovery checks appropriate to the interface; and
- support channel, change notice, service status, audit evidence, and offboarding procedures.
A documented control, a visible setting, and an enforced behavior are three different evidence points. When feasible, we test whether a saved administrator setting changes a synthetic user’s experience. We do not test controls with real student accounts or records.
District deployability review
Deployability asks whether the named configuration can operate inside a district’s real constraints. The review defines assumptions for:
- identity source, account lifecycle, role mapping, and delegated administration;
- managed and unmanaged devices, supported operating systems, browser or native application requirements, and network dependencies;
- data origin, destination, residency, retention, deletion, backup, and district export requirements;
- accessibility, language, family communication, help desk, training, and classroom-change workload;
- licensing, metered usage, local hardware, staffing, monitoring, and foreseeable support cost;
- implementation sequence, pilot boundary, rollback, outage behavior, incident response, and vendor exit; and
- the district policy, DPA, procurement, security, curriculum, and legal decisions still required.
The result is a dated fit assessment against declared district requirements, not a claim that the product is deployable everywhere. A local model and a hosted education product can both be viable while imposing very different hardware, staffing, data, and guardrail responsibilities.
The repeatable evaluation protocol
1. Freeze the cases and scoring rules
Before running any named configuration, we version the scenario packet, expected content, unacceptable shortcuts, applicable rubric dimensions, operator script, and analysis plan. Sixteen case descriptions are public. Eight cases are held out to reduce direct optimization against the complete suite. We publish the held-out categories and scoring method, but not their exact prompts, source materials, or answer keys until the cases are retired and replaced.
No case uses a real learner profile. All student work, misconceptions, records, and accommodation descriptions are synthetic. Source materials are public domain, licensed for the test, or created for the evaluation.
2. Run three independent repeats
Each configuration receives three independent runs per scenario, producing 72 conversations across the 24-scenario suite. Every repeat begins in a clean new conversation. Memory, personalization, search, tools, and retrieval follow the frozen scenario configuration rather than the operator’s preference.
Run order is randomized. Fixed multi-turn cases use the same prewritten student turns. If a case requires a branch, the packet defines the allowed branches in advance, and a non-rater operator logs which branch was selected. Operators do not coach a weak response into a passing one.
Three repeats reveal obvious variability without pretending to characterize the full output distribution. We report all three outcomes, the median where a numerical summary is appropriate, and the observed range. We do not report a precision statistic that the sample cannot support.
3. Preserve raw outputs and blind the review set
The output is preserved exactly. We may remove product names, interface chrome, timestamps, and other identity cues for blinding, but we do not rewrite grammar, citations, refusals, or substantive content. Each transcript receives a random code, and transcript order is randomized independently for each rater.
The identity key remains unavailable to raters until scoring and adjudication are complete. A failure to fully blind a product, such as a distinctive self-reference inside the response, is recorded as a limitation.
4. Use two educator raters and adjudication
Two qualified educator raters independently score every transcript. For this protocol, a qualified rater is a current or former K-12 educator, curriculum specialist, or instructional leader with relevant grade-band or subject experience. Each scorecard discloses the raters’ applicable experience and potential conflicts. Raters receive the learning objective, synthetic learner state, content answer key, rubric, and case-specific boundaries. They do not receive the product identity or the other rater’s score.
Before scoring the evaluation set, both raters independently score a six-transcript calibration set that is excluded from the published product results. They compare their use of the rubric anchors, resolve interpretation differences, and repeat any failed calibration dimension before the scored review begins. Calibration does not authorize changing the rubric after product identities are revealed.
An educator adjudicator reviews a case when:
- raters differ by two or more points on any 0 to 3 dimension;
- one rater records a critical factual error and the other does not;
- one rater records an answer-giving or integrity failure and the other does not; or
- the raters disagree about whether the response satisfied the scenario’s core instructional boundary.
The scorecard reports exact raw agreement and linearly weighted Cohen’s kappa for each ordinal dimension when the statistic is estimable. If it is not estimable, the report says why rather than substituting a different statistic after seeing the results. It also reports how many transcripts required adjudication. The adjudicated record preserves both original ratings rather than erasing disagreement.
5. Separate automated checks from educator judgment
Automated checks can verify exact answers, required citations, duplicate items, reading level, formatting, latency, token use, or whether a direct answer appeared where it was prohibited. They cannot replace educator review of explanation quality, cognitive demand, scaffolding, or learner agency.
We publish which fields were machine checked, which were educator rated, and which were derived. If an AI system assists with coding or analysis, that use is disclosed and a person verifies the result.
Initial 24-scenario K-12 design
The first suite covers student tutoring, feedback, information literacy, teacher workload, accessibility, adaptation, and academic-integrity boundaries. Public cases expose the task structure. Held-out cases disclose the purpose and level while reserving the exact prompt and answer key.
| ID | Disclosure | Role and domain | Defined task and success boundary |
|---|---|---|---|
| P01 | Public | Grade 5 student, mathematics | Diagnose an invented fraction-denominator misconception and provide one next-step hint without completing the problem |
| P02 | Public | Grade 8 student, algebra | Use the student’s incorrect equation work to ask a targeted diagnostic question before explaining |
| P03 | Public | Grade 8 student, physical science | Correct a plausible false premise about force and motion with an age-appropriate explanation and a check for understanding |
| P04 | Public | Grade 10 student, biology | Help distinguish correlation from causation in a synthetic experiment without inventing evidence |
| P05 | Public | Grade 6 student, writing | Give rubric-linked feedback on an invented paragraph while preserving the student’s words and not rewriting it |
| P06 | Public | Grade 10 student, argument writing | Identify the weakest evidence-to-claim connection and prompt the learner to revise it |
| P07 | Public | Grade 8 student, civics | Compare two provided sources, separate fact from interpretation, and request corroboration where needed |
| P08 | Public | Grade 11 student, U.S. history | Evaluate whether a claim is supported by a supplied primary-source excerpt and acknowledge evidentiary limits |
| P09 | Public | Grade 7 student, mathematics homework | Resist a direct-answer request and offer a graduated hint that preserves productive effort |
| P10 | Public | Grade 10 student, computer science | Diagnose a bug in short synthetic code through questions and a minimal hint rather than replacing the solution |
| P11 | Public | Grade 4 teacher, mathematics | Generate 30 standards-linked exit-ticket variants with keyed answers, no duplicate stems, and reviewable difficulty labels |
| P12 | Public | Grade 8 teacher, writing | Apply a supplied rubric to ten synthetic responses and identify the minimum teacher corrections required |
| P13 | Public | Grade 10 teacher, science | Adapt one lesson explanation to two reading-support levels without reducing the scientific learning objective |
| P14 | Public | Grade 5 teacher, family communication | Produce a plain-language multilingual family explanation while preserving dates, requirements, and key academic terms |
| P15 | Public | Grade 8 teacher, accessibility | Reformat a synthetic handout for screen-reader navigation without losing content, labels, or sequence |
| P16 | Public | Grade 9 student, multi-turn tutoring | Adapt a second hint after the learner reveals a different misconception, then ask for self-explanation |
| H17 | Held out | Grade 6 student, mathematics | Diagnose a novel misconception from unseen work and preserve the target cognitive demand |
| H18 | Held out | Grade 11 student, science | Interpret an unseen data display, calibrate uncertainty, and distinguish observation from inference |
| H19 | Held out | Grade 7 student, writing | Maintain feedback-only behavior after the learner pressures the tutor to rewrite the full response |
| H20 | Held out | Grade 12 student, civics | Synthesize competing sources without treating source confidence as factual certainty |
| H21 | Held out | Grade 9 student, metacognition | Elicit a self-explanation, detect a shallow explanation, and prompt a transfer-oriented revision |
| H22 | Held out | Middle school teacher, assessment | Complete a repetitive assessment-generation task with verified keys, variation, and measurable correction burden |
| H23 | Held out | Student support, accessibility | Apply a synthetic accommodation without disclosing private records or lowering the learning objective |
| H24 | Held out | Secondary student, integrity boundary | Maintain a teacher-defined no-direct-completion rule across a realistic multi-turn attempt to obtain the answer |
The suite is deliberately not an exam leaderboard. K-12EduBench provides a much larger subject-knowledge and problem-solving benchmark. Our 24 scenarios instead sample district-facing applications and instructional behavior. A future scorecard may report K-12EduBench or another validated benchmark separately, but it will not substitute benchmark accuracy for classroom suitability.
The pedagogy rubric
The rubric adapts research themes including active learning, cognitive-load management, metacognition, curiosity, adaptivity, subject knowledge, problem-solving process, and educational-goal cognition. It adds explicit K-12 boundaries for factual calibration and preserving student work.
Each applicable dimension receives 0 to 3 points. Not applicable is excluded from the denominator. Case packets declare applicable dimensions before testing so a reviewer cannot choose a favorable denominator after seeing the output.
| Dimension | 0: harmful or absent | 1: weak | 2: adequate | 3: strong |
|---|---|---|---|---|
| Learning-goal alignment | Replaces or contradicts the intended learning | Addresses the topic but misses the cognitive goal | Supports the stated goal with minor drift | Keeps every move focused on the goal and target cognitive demand |
| Correctness and calibration | Contains a material error or fabricated support | Mixes useful content with uncorrected error or unjustified certainty | Is substantively correct and marks important uncertainty | Is correct, evidence-aware, and helps the learner verify limits or sources |
| Diagnosis and relevance | Ignores the learner’s work or misconception | Gives generic advice with little diagnosis | Uses available evidence to target the main need | Tests its diagnosis, distinguishes likely misconceptions, and responds precisely |
| Active learning and learner agency | Supplies prohibited completion or removes meaningful learner work | Asks superficial questions after doing most of the work | Uses questions, hints, or partial steps that preserve learner action | Sequences productive effort and gives the learner a meaningful next decision |
| Scaffolding and cognitive load | Overwhelms, withholds essential support, or creates confusion | Provides poorly sequenced or mismatched support | Breaks the task into manageable steps with suitable detail | Adjusts support deliberately, fades help, and connects steps into a coherent model |
| Adaptivity and accessibility | Stereotypes, exposes synthetic private context, or lowers the objective unnecessarily | Changes surface wording without addressing the stated need | Adapts format, language, or support while preserving the objective | Confirms the need, offers an effective accessible path, and preserves rigor and dignity |
| Metacognition and transfer | Encourages copying or confidence without reflection | Adds a generic “check your work” prompt | Elicits explanation, monitoring, or a related application | Helps the learner explain strategy, test confidence, and transfer the idea to a novel case |
| Instructional and integrity boundary | Violates the teacher-defined boundary or gives a disguised answer | Initially violates or inconsistently follows the boundary | Maintains the boundary and offers a useful allowed alternative | Maintains it across turns, explains the constraint briefly, and redirects to productive learning |
Critical factual errors and prohibited direct completion are reported as failure rates in addition to rubric scores. A high average cannot hide a small number of severe failures.
Scorecards report, at minimum:
- the distribution of 0, 1, 2, and 3 ratings by dimension;
- the three-repeat pass count for every scenario;
- critical factual-error, unsupported-citation, and boundary-failure counts;
- rater agreement and adjudication rate;
- public-case and held-out-case results separately; and
- qualitative examples selected by a rule declared before model identities are revealed.
Operational performance and cost
Operational tests use the same frozen task inputs, but they answer a different question from pedagogy. We record:
- time to first visible output and total completion time;
- successful completion, provider error, timeout, truncation, and operator retry;
- required user turns and whether the product completed the requested format;
- input and output tokens when the endpoint exposes them;
- dated API cost using the provider’s published units, without pretending a consumer subscription has a per-task price;
- for local models, measured runtime and hardware configuration, with energy or ownership cost reported only when directly instrumented and defined;
- valid items per minute and valid items per dollar for repetitive generation tasks; and
- educator correction count and correction time needed to make the output usable.
With three repeats, we publish the median and full range rather than a fragile percentile. Product UI, API, and local-runtime measurements remain in separate tables because their prices, infrastructure, and control layers are not equivalent.
Provider claims and observed results use different labels
Every fact in a scorecard is labeled by evidence type:
- Provider documented: stated in official terms, privacy materials, technical documentation, or account guidance as of the review date.
- Contract confirmed: present in the evaluator’s applicable signed terms or directly confirmed for that tested account. A district still needs its own contract review.
- Observed in configuration: directly visible or reproducible in the named test account or runtime.
- Measured: produced by the frozen task and measurement protocol.
- Educator rated: a blinded rubric judgment, with agreement and adjudication disclosed.
- Not disclosed or not testable: evidence was unavailable. We do not convert missing evidence into a favorable assumption or an automatic zero.
Provider documentation cannot override an observed failure. One observed run also cannot disprove every deployment covered by broad provider documentation. Both remain visible, and the scope of each statement stays narrow.
For the wider governance context, see how school districts can govern student use of AI tools and why data loss prevention matters in K-12 AI.
Pedagogical behavior is not evidence that students learned
Pedagogical behavior is not evidence that students learned.
A model can ask thoughtful questions, give accurate hints, and receive strong educator ratings without improving retention or transfer for a real learner.
This boundary matters because assisted task performance can move in the opposite direction from unassisted learning. In the PNAS high school mathematics field experiment cited below, access to a general GPT-4-style tutor improved practice performance, yet the group using the less guarded interface performed worse than the control group when AI access was removed. The guarded tutor mitigated that negative effect but did not establish a positive unassisted learning gain in that study. The finding supports testing the instructional harness around a model, not assuming that fluent assistance produces learning.
We will use learning language only when a separate study includes:
- a preregistered research question, outcomes, exclusion rules, and analysis plan;
- an appropriate comparison condition and assignment method, with baseline equivalence addressed;
- a pretest measuring the targeted knowledge or skill before the intervention;
- a defined implementation period and fidelity measures showing what participants actually received;
- an AI-assisted practice measure reported separately from learning outcomes;
- an unassisted post-test using suitable parallel items;
- a delayed transfer or retention assessment requiring the skill without AI on meaningfully new items;
- reporting of sample, attrition, missing data, uncertainty, effect sizes, subgroup analysis rules, and limitations; and
- district authorization plus any required institutional review, consent, assent, privacy, and human-subject protections.
The What Works Clearinghouse standards inform how we assess comparison designs, baseline equivalence, attrition, outcome validity, and confounding. A classroom pilot intended for product feedback is not relabeled as a learning study after the fact.
Uncertainty, changes, corrections, and reruns
AI products change quickly, and a score is not permanent. Every scorecard includes:
- a publication date and evaluation window;
- methodology, scenario-set, and rubric version;
- exact tested configurations;
- known sources of uncertainty and non-generalizability;
- raw counts alongside summaries;
- a dated change log;
- correction and provider-response notes; and
- a superseded label when a material rerun replaces an earlier result.
A material change includes a new model snapshot, hidden or automatic routing change, changed system behavior, new search or tool access, altered account controls, a new retention or training-use term, a safety-policy change, a scoring-rule change, or a scenario leak. We rerun the affected scope and never silently overwrite the earlier dated result.
Results generalize only to the tested cases and configuration. Three repeats do not reveal every rare failure. Two educator raters do not represent every school community. Held-out cases reduce direct optimization but do not eliminate benchmark contamination. These limits appear beside the results, not in fine print.
Frequently asked questions
What does a TrueMadeAI K-12 AI scorecard measure?
It reports five separate evidence layers: product and account controls, district deployability, operational performance, pedagogical behavior, and, only when supported by an appropriately designed study, actual learning outcomes. The layers are not collapsed into one universal score.
Does the highest pedagogy score identify the best AI product for every school?
No. A product can perform well on selected tutoring scenarios while being unsuitable for a district’s age groups, data boundaries, account controls, accessibility needs, deployment model, or instructional purpose. Results apply only to the tested configuration and cases.
Is strong pedagogical behavior evidence that students learned?
No. A transcript can show behavior consistent with good tutoring, but actual learning claims require student outcome evidence such as a pretest, an unassisted post-test, a delayed transfer measure, an appropriate comparison, and required research approvals.
Why test a product interface separately from its API model?
A product interface can add system instructions, account policies, browsing, tools, memory, safety controls, retrieval, and model routing. An API test can isolate a named model configuration more precisely, but it does not reproduce the complete student or educator product experience.
How do you prevent private student information from entering evaluations?
The scorecard suite uses synthetic personas, invented student work, public-domain or licensed materials, and fabricated records. No real student records, names, transcripts, accommodations, or education records are used.
How often are AI products retested?
A dated result remains tied to its recorded configuration. We rerun affected cases after a material model, interface, policy, account, tool, retrieval, or safety change and record the reason in the scorecard change log.
Can providers review a scorecard before publication?
A provider may identify a factual configuration error or supply official documentation, but it cannot alter observed outputs or educator ratings. Corrections, responses, and reruns are labeled so readers can distinguish them from the original result.
Sources
- K-12EduBench: A Benchmark for Evaluating Large Language Models’ Knowledge, Problem-Solving, and Educational Goal Cognition in K-12 Education, AAAI
- LearnLM: Improving Gemini for Learning, LearnLM Team
- Evaluating Gemini in an Arena for Learning, LearnLM Team
- Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics, PNAS
- AI Risk Management Framework Core, NIST
- Generative AI Profile, NIST
- What Works Clearinghouse Procedures and Standards Handbook, Version 5.0, U.S. Department of Education
This methodology is an evaluation framework, not a certification, endorsement, legal opinion, or guarantee of safety, privacy, accessibility, instructional quality, or learning impact.