Current status: This is a practical evaluation framework, not legal advice or a compliance determination. Districts should apply their own policies, contracts, state requirements, and counsel.
A school district should pilot an AI tool as a bounded test of a specific use, account configuration, user group, data boundary, and operating process. Set success and stop criteria before testing. Begin with synthetic data. Test real workflows and failure modes. Document the results, then approve only the configuration the district actually evaluated.
The rest of the work is making that boundary real.
The common mistake is to run an impressive demo, collect a few enthusiastic comments, and call the product approved. A district-wide rollout is not a larger demo. It is an operating decision involving students, staff, data, instruction, support, and accountability.
This guide turns that decision into a practical school district AI pilot plan.
Pilot a use, not a vendor name
“We piloted Gemini” or “we approved ChatGPT” is not specific enough.
One AI product can contain different models, account types, age restrictions, privacy terms, connected services, file paths, administrator controls, and classroom uses. A configuration that is reasonable for staff brainstorming may not be appropriate for a student asking for writing feedback. A district-built family assistant using public documents has a different data boundary from an authenticated staff workflow.
Define the pilot as a complete sentence:
We are testing [specific product and configuration] for [defined task] with [named users], using [permitted data], on [managed environment], under [approved controls], so that [decision owner] can decide [what may happen next].
If the team cannot complete that sentence, the pilot is not ready.
The AI tool vetting and approval template provides the full approval record. The pilot plan below focuses on how to generate the evidence that belongs in that record.
Name the decision before the test begins
A pilot should answer a decision, not merely create activity.
Examples include:
- Can high school teachers use an approved AI workspace for lesson planning without entering student records?
- Can students use a supported AI product for guided brainstorming under grade and teacher rules?
- Can a public district assistant answer calendar and handbook questions with inspectable citations?
- Can an internal IT assistant retrieve only from an approved technical knowledge collection?
Write down the possible outcomes before anyone starts testing:
- Go: the tested configuration may move to a defined next cohort.
- Conditional go: it may proceed only after named controls or contract terms are in place.
- Revise and retest: the use is promising, but the current configuration did not satisfy the criteria.
- No go: the district will not deploy this use under the tested conditions.
A “go” is not permanent approval of every future model, feature, or use. It is permission for the scope the district tested.
Build the pilot team and assign one owner
A small pilot still crosses district functions. Include the people necessary to make and operate the decision:
- an executive sponsor;
- a pilot owner with authority to pause the test;
- technology and security;
- student privacy or data governance;
- curriculum and instruction when learning is in scope;
- accessibility and multilingual support;
- procurement or legal review when contracts are involved;
- a school leader and representative educators or staff; and
- the help desk or team that will support actual users.
Do not let responsibility dissolve into a committee. One person should own the pilot record, open issues, decision meeting, and exit plan.
Use a phased school district AI pilot plan
The phases matter more than an arbitrary number of days. A simple assistant may move through them quickly. A student-facing tool with roster data, integrations, or consequential workflows should take longer.
Phase 1: define scope and data boundaries
Record the exact product, model or model family where visible, account type, administrator settings, browser or application surface, users, purpose, and prohibited uses.
Map each data element the pilot could receive:
| Data category | Example | Pilot rule |
|---|---|---|
| Public district information | Calendar, handbook, public policy | May be allowed for a defined public-information use |
| Operational staff content | Draft agenda, generic procedure | Review purpose, access, retention, and provider terms |
| Student directory information | District-defined directory fields | Do not assume it is unrestricted; apply district policy |
| Student records or sensitive data | Grades, discipline, disability, counseling | Exclude until a specific lawful and approved data flow exists |
| Credentials and secrets | Passwords, tokens, private keys | Never enter into the AI workflow |
The K-12 AI data-boundary framework helps teams distinguish authorization, minimization, retention, DLP, and evidence. These are related controls, not interchangeable promises.
The U.S. Department of Education’s student privacy resources advise schools to understand how an online service collects, uses, and transmits information before deciding whether to use it. That review belongs before real student data enters a pilot, not after.
Phase 2: test with synthetic and nonstudent data
Build realistic test cases without using real student records. Use invented names, fictional rosters, synthetic documents, sample prompts, and public district information.
Test what should work and what should fail:
- an ordinary approved request;
- an out-of-scope request;
- an attempt to enter restricted data;
- an inappropriate or harmful request;
- an unsupported file or content type;
- a user with the wrong role or account;
- a missing citation or weak source;
- an unavailable model, policy service, or integration;
- an administrator changing a setting; and
- a request to disable or remove the tool.
Record the exact configuration and date. AI systems and vendor interfaces change. A result without a configuration record is difficult to reproduce and easy to overstate.
NIST’s AI Risk Management Framework calls for testing before deployment and monitoring while systems are operating. It also emphasizes documenting test sets, metrics, methods, and performance under conditions similar to the intended deployment. A school pilot can apply that principle without turning the project into a research lab.
Phase 3: run a limited, representative cohort
After the district approves the data flow and initial controls, move to a small cohort that represents the real environment.
Choose participants deliberately. Include different schools, roles, grade bands, devices, accessibility needs, and support conditions when they are relevant to the intended use. Do not select only the most enthusiastic early adopters and then treat their experience as proof for the entire district.
Give participants:
- the approved purpose;
- examples of permitted and prohibited data;
- instructions for reporting an incorrect or concerning result;
- a clear human support path;
- notice of what operational evidence is collected; and
- the date and condition under which the pilot ends.
For instructional pilots, evaluate more than whether users liked the tool. The district should distinguish product usability, model performance, instructional fit, and evidence of student learning. Our K-12 AI evaluation methodology explains why a convincing answer is not, by itself, proof that learning improved.
Phase 4: monitor operations, not only outputs
The district needs to learn what operating the tool actually requires.
Track:
- support questions and time to resolution;
- incorrect, unsafe, or out-of-scope responses;
- data-handling events and false positives;
- accessibility and language barriers;
- account, identity, and permission failures;
- changes to models, terms, interfaces, or administrator controls;
- uptime and dependency failures;
- user bypasses and unapproved alternatives; and
- whether staff can explain and apply the rules consistently.
This is where many attractive pilots fall apart. The product may perform its headline task while creating an operating burden the district did not budget for.
Phase 5: make and publish the decision internally
At the end of the pilot, compare the evidence with the criteria set at the beginning. Record the outcome, decision owners, conditions, known limitations, review date, and rollback path.
If the pilot advances, tell users exactly what was approved. A concise approval notice should identify:
- the permitted use;
- eligible users and accounts;
- allowed and prohibited data;
- required settings and controls;
- where to get support;
- how to report a problem; and
- when the district will review the approval again.
An approval that nobody can explain will not govern real behavior.
Set success criteria and stop rules before launch
Success criteria should be observable. Stop rules should be specific enough that the pilot owner can act without convening a philosophical debate during an incident.
| Area | Example success criterion | Example stop or pause rule |
|---|---|---|
| Purpose | Representative users complete the approved workflow | The tool repeatedly substitutes an unapproved task |
| Data | Test data stays within the approved flow and settings | Restricted data is exposed to an unapproved recipient or path |
| Access | Only the intended cohort can use the configured experience | An unauthorized role reaches a protected function or source |
| Quality | Results meet the documented review threshold for the use | Errors create a material instructional, safety, or operational risk |
| Safety | Refusal and escalation paths work in the tested scenarios | A high-severity scenario bypasses the expected control |
| Accessibility | The workflow is usable with required assistive technology | A required user group cannot access the core workflow |
| Operations | Support owners can diagnose, disable, and recover the service | The district cannot reliably pause the tool or identify its configuration |
| Change control | Material model and product changes trigger review | The provider changes a load-bearing feature without an acceptable review path |
These are examples, not universal thresholds. Each district should define severity, evidence, and decision authority for its own use.
Keep a pilot evidence pack
A useful pilot leaves behind more than a slide deck.
Keep one durable record containing:
- the use-case statement and decision owner;
- product, model, account, settings, and version information available at the time;
- audience, cohort, duration, and deployment environment;
- approved and prohibited data;
- data-flow and access diagrams;
- relevant agreements, terms, and district approvals;
- test scenarios, expected results, and observed results;
- incidents, false positives, support issues, and unresolved limitations;
- accessibility, privacy, security, and instructional review notes;
- success criteria, stop rules, and final decision;
- conditions for broader use; and
- review, renewal, rollback, and decommission dates.
The K-12 AI governance software buyer’s guide includes a broader RFP and validation checklist. Use it when the pilot is also comparing governance architectures or vendors.
What not to do during an AI pilot
Avoid these shortcuts:
- Do not approve a logo. Approve a defined use and configuration.
- Do not use real student records to prove a privacy control works. Start with synthetic data.
- Do not test only ideal prompts. Test mistakes, misuse, ambiguity, outages, and attempts to exceed scope.
- Do not rely only on vendor demonstrations. District staff should run the actual scenarios.
- Do not confuse a DPA with runtime authorization. A contract and an operational control solve different problems.
- Do not measure only sentiment. Enthusiasm matters, but so do accuracy, support, accessibility, data handling, and instructional fit.
- Do not omit the exit. The district should know how to disable access, preserve necessary records, notify users, and remove data when applicable.
- Do not let a pilot become permanent by inertia. End with an explicit decision and review date.
How Tenet can support a bounded AI pilot
Tenet separates two kinds of district AI activity.
Tenet Edge supports governed direct use on district-managed Chrome across supported AI products and surfaces. Available controls vary by product, content type, configuration, and rollout scope. The supported-products capability matrix states those boundaries publicly.
Tenet Gateway is the founding-district program for governing approved backend AI operations inside district applications. It is designed around application identity, purpose, permitted data, eligible model deployments, budget, and content-minimized decision evidence before an approved model is called.
Neither product eliminates district review, contracts, professional judgment, or application-specific testing. The value is making configured rules and evidence part of the operating path instead of leaving the pilot as a one-time document.
If your district is preparing to test AI, request a Tenet pilot or begin with the AI governance readiness assessment.
Sources
- AI Risk Management Framework Core, National Institute of Standards and Technology
- NIST AI RMF Playbook, National Institute of Standards and Technology
- Privacy and Education Technology, U.S. Department of Education Student Privacy Policy Office
- Protecting Student Privacy While Using Online Educational Services, U.S. Department of Education Student Privacy Policy Office
This guide is practical district planning material. It does not determine whether a particular product, contract, data flow, or use complies with law or district policy.