AI & Privacy

Is your student data
training AI models?

One of the most important privacy questions of the moment: is the student data your SIS vendor holds being used to train AI models? Most districts don't know.

Talk to Alma
The short answer

Student data should not be used to train external AI models. Contract language should explicitly prohibit vendor use of your data for training foundation models or improving AI features that will be sold to other customers. If your vendor uses AI features that need training data, those should be trained on de‑identified or synthetic data, with opt‑in consent from your district. Ask the question directly, in writing, and require the answer in the contract.

Why this question matters

Student data is uniquely valuable, and uniquely at risk.

Once student data has been used to train an AI model, it cannot be extracted. The model retains patterns from that data indefinitely. Districts that don't ask this question now may find their students' data embedded in commercial AI models they never consented to.

What to verify with your SIS vendor:

  • No training on district data: Explicit contract language prohibiting use of your data to train AI models
  • Third-party model access: If the vendor uses external LLMs (OpenAI, Anthropic, Google), what happens to the data sent to those models?
  • Retention policies: How long are AI queries and outputs retained?
  • De‑identification standards: If data is de‑identified for training, what standard is used?
  • Opt‑in versus opt‑out: Are AI features opt‑in for districts?
  • Generative AI features: Special attention to any features that generate text, feedback, or recommendations
  • Prompt retention: Do teacher or administrator prompts get stored, and if so where?

Common questions, answered directly.

The honest answer is: many districts don't know. Contract language from a few years ago rarely addressed AI training explicitly. Ask your current vendor directly, in writing, whether any of your district's data has been used to train AI models, whether shared with third-party AI providers, or used in any way beyond delivering the SIS service to your district.

FERPA restricts disclosure of personally identifiable information from student education records without consent. Using student data to train AI models arguably constitutes a use beyond the original educational purpose and may require additional consent. FERPA guidance from the Department of Education on AI training is still emerging, but the conservative interpretation is that vendors should not train models on identifiable student data without explicit district authorization.

A clear statement that your district's data will not be used to train AI models, develop AI products for other customers, or be shared with third parties for AI training purposes. Language should explicitly cover both direct training and inclusion in derived datasets. Include this in the data processing agreement, not just the general terms of service.

Ask Alma to provide this language in writing as part of your data processing agreement before signing. A vendor willing to put a no-training commitment into contract language, rather than just a sales conversation, is giving you something enforceable.

De‑identification standards vary. Aggregated statistics (like average attendance rates by grade level) are generally safe for use. Individual student data with names, IDs, and specific patterns removed may still be re-identifiable if combined with other data. Ask your vendor to specify their de‑identification standard, whether it meets FERPA de‑identification requirements, and who reviews it.

Districts should be able to opt out. Ask specifically whether AI features are enabled by default and whether disabling them affects other platform functionality. Some vendors have started making AI opt‑in for districts. Others still enable AI by default and require opt‑out. The former is the standard districts should demand.

Pay special attention. Features that let teachers or administrators generate text (feedback, communications, lesson plans) using AI often send prompts and context to external LLM providers. Ask whether those prompts include student data, whether they're retained by the LLM provider, and whether the LLM provider uses them for training. OpenAI's terms for API use are different from consumer ChatGPT terms, and vendors should be able to explain which they use.

That's a reasonable district-level policy while your team gets clarity. Some districts have implemented AI moratoriums until they can review each feature. Others have created AI review committees that must approve any AI‑enabled feature before it's used with student data. The right approach depends on your district's risk tolerance and technical capacity.

Alma treats student data as belonging to the district, not the vendor. Ask directly about AI features, how they work, what data they access, and how that data is handled. Straight answers about specific practices, not vague reassurances, are what to look for from any vendor.

Bring this specific question to any SIS vendor evaluation: is our student data used to train AI models, and can that prohibition be written into our contract?

Ready to see how Alma handles this?

Get straight answers to the questions this guide raises. Schedule a walkthrough with Alma.

Schedule a Demo