Personalized Learning Data Privacy: Who Owns Student Data?
“She doesn’t let me finish my sentences, she doesn’t listen,” a first-grade girl told her mother after coming home from school. The complaint wasn’t about a classmate. It wasn’t about a teacher. It was about Amira, an artificial intelligence that had been listening to her read aloud in class, recording her voice, analysing her every stumble and pause, and judging her without ever letting her finish.
I read that story in a Christian Science Monitor investigation and could not stop thinking about it. The mother’s name is Cassie Fleming. Her daughter is in first grade. And the AI reading tool her school deployed, without ever asking her permission, has now captured hundreds of thousands of voice recordings of New Mexico schoolchildren. Young kids. Reading aloud into a screen while an algorithm listened back, judged them, and stored their voices somewhere those children and their parents will never see.
That detail, a six-year-old frustrated that an AI wouldn’t let her finish her sentences, captures everything that is both promising and troubling about the collision of artificial intelligence and education. The technology can listen. It can analyse. It can personalise. But should a child’s voice, literally, their voice, become part of a dataset they will never see, stored somewhere they will never know, governed by policies their parents may never have truly understood?
This is not just about one reading app. It is about a fundamental question that education systems around the world are not asking loudly enough: are we trading our children’s personal data for personalised learning without knowing the full cost?
⚡
Key Takeaways
Hundreds of thousands of voice recordings
AI reading tools like Amira capture children’s actual voices, a form of biometric data that most school privacy policies never contemplated.
Parents were never fully informed
The CS Monitor investigation found that many New Mexico parents had no idea their children’s voices were being recorded and stored.
96% of edtech apps share student data
Research shows the vast majority of educational technology tools share student data with third parties, often without clear disclosure.
Privacy laws are decades behind
FERPA was written in 1974. COPPA in 1998. Neither was designed for AI systems that capture biometric and behavioral data in real time.
What Actually Happens When a Child Reads to Amira?
Let me paint the picture. A first-grade student sits in front of a tablet or laptop. The screen shows a passage. The child reads aloud. Amira listens, not passively, but actively. The AI analyzes every syllable, every pause, every mispronunciation. It tracks which words the child stumbles on, how long they hesitate, and whether they self-correct.
According to Amira Learning’s own product description, the system “continuously and authentically listens as students read aloud” and “assesses proficiency.” This means the AI isn’t just grading a one-time test, it is building a running profile of the child’s reading ability over weeks and months.
Now, Amira is not the only tool doing this. DreamBox, i-Ready, Khan Academy’s Khanmigo, every major adaptive learning platform collects granular data on how students think, learn, and perform. This is part of a broader pattern I explored when I wrote about the hidden risks of AI in education, the data collection infrastructure runs deeper than most people realize. But Amira is different in one crucial way: it records children’s voices.
Voice is not just data. Voice is biometric. It can reveal a child’s age, gender, emotional state, and, with enough processing, things we haven’t even thought of yet. Under laws like Illinois’s BIPA (Biometric Information Privacy Act), voice recordings are treated with heightened protection. Yet in education, we are handing over children’s voice data with far less scrutiny than we apply to a fingerprint scanner in a gym locker room.

Where Does All This Data Actually Go?
This is where the trail gets murky, and where the NBC investigation raised the most uncomfortable questions.
When Amira records a child reading in a New Mexico classroom, that audio file travels from the school’s device to Amira Learning’s cloud infrastructure. Amira Learning, which is now part of Istation (a larger edtech company), processes those recordings to generate reading assessment reports for teachers.
But here are the questions that parents in New Mexico, and parents everywhere, should be asking:
- 🔹 Storage location: Where exactly are these voice recordings stored? On servers in the United States? Overseas? Who physically controls those servers?
- 🔹 Retention period: How long are the recordings kept? Does a child’s first-grade reading sample follow them? If a child leaves the school or moves states, what happens to their voice data?
- 🔹 Third-party access: Are the recordings shared with subcontractors, cloud providers, or analytics partners? Does Amira Learning’s privacy policy allow data to be shared for “product improvement”, a phrase that can cover almost anything?
- 🔹 AI training: Can the voice data be used to train future AI models? This is a critical distinction. Is the data used only to help that specific child, or does it become training material for the next version of the product, or even for entirely different AI systems?
- 🔹 Deletion rights: If a parent wants their child’s voice recordings deleted, is that possible? And can they verify deletion actually happened?
The uncomfortable truth is that for most edtech tools deployed in schools today, clear, public answers to these questions do not exist, or are buried so deep in privacy policies that no reasonable person would ever find them.
The Scale of the Problem: How Much Student Data Is Really Out There?
The Amira case is alarming on its own. But it is part of a much larger pattern.
In 2022, a landmark study by Internet Safety Labs examined 1,357 edtech applications used in K-12 schools across the United States. The findings were staggering: 96% of these apps shared student data with third parties. Not 20%. Not half. Ninety-six percent.
Another major study from the Center for Democracy and Technology (CDT) found that the average school district uses over 1,400 different edtech tools each month, and a large proportion of those tools had never undergone a formal privacy review by the district.
Meanwhile, a 2025 report by the Future of Privacy Forum (FPF) on algorithmic personalization in youth online experiences warned that data-driven personalization in education, while offering real benefits, is creating detailed digital profiles of children that may have consequences far beyond the classroom.
Human Rights Watch, in a major 2022 report titled “How Dare They Peep into My Private Life?” documented how children’s data collected by edtech platforms was being used to make automated decisions about their academic trajectories, with little transparency or recourse for families who disagreed.
This isn’t a niche worry. It’s a systemic issue affecting millions of children, across thousands of schools, every single day.
| Privacy Concern | What Parents Should Ask | What Schools Should Verify |
|---|---|---|
| Voice/biometric data | Is my child’s voice being recorded and stored? | Where are audio files stored and for how long? |
| Third-party sharing | Which companies have access to my child’s data? | Do vendor contracts explicitly prohibit data resale? |
| AI model training | Can recordings be used to train future AI systems? | Does the vendor’s policy separate service data from training data? |
| Data retention | What happens to data when my child leaves the school? | Is there a guaranteed deletion process with verification? |
| Parental consent | Was I given meaningful notice and a real choice? | Is opt-in consent obtained before data collection begins? |
| FERPA/COPPA compliance | Does this tool comply with federal student privacy laws? | Has the district’s data protection officer reviewed this tool? |
The Law Is Decades Behind the Technology
To understand why this is happening, you need to understand the legal framework, and how outdated it really is.
The main federal law protecting student data in the United States is FERPA, the Family Educational Rights and Privacy Act. FERPA was passed in 1974. For context, that is the same year the Rubik’s Cube was invented. It was written for a world of paper files in filing cabinets, not cloud-based AI systems that record children’s voices.
FERPA protects “education records”, but the definition was built for report cards, disciplinary files, and attendance records. When an AI captures voice recordings, reading patterns, eye-tracking data, or emotional state analysis, is that an “education record”? The law doesn’t clearly say.
Then there is COPPA, the Children’s Online Privacy Protection Act, passed in 1998. COPPA requires parental consent before collecting personal information from children under 13. But in schools, a loophole allows schools to consent on behalf of parents when the tool is used for educational purposes. This means a parent may never even see the privacy policy for a tool their child uses every day.
The European Union’s GDPR offers stronger protections, including the “right to be forgotten” and stricter rules on biometric data, but enforcement in the edtech sector has been inconsistent.
As the Future of Privacy Forum noted in their 2026 report, “the legal frameworks governing student data were not designed for AI systems that continuously collect, analyze, and act upon behavioral data in real time.” That is a polite way of saying: we are flying blind.
InBloom: The Warning We Should Have Heeded
The Amira controversy is not the first time America has faced this question. It is the latest chapter in a story that started over a decade ago.
In 2013, a nonprofit called inBloom launched with $100 million in funding from the Gates Foundation and the Carnegie Corporation. The idea was ambitious: create a centralized data platform where schools could store and manage student information, making it easier for edtech tools to personalize learning.
But when parents learned what kind of data would be stored, including discipline records, disability status, family relationships, and even bus stop locations, the backlash was swift and fierce. Parent protests erupted in New York, Louisiana, and several other states. Within a year, inBloom shut down entirely.
The inBloom case showed that when parents do learn what data is being collected, they often say no. But the lesson the edtech industry seems to have drawn is not “collect less data”, it is “don’t tell parents until after the tool is already in classrooms.”
Amira was already deployed across New Mexico before many parents knew it existed. By the time Cassie Fleming’s daughter came home complaining about the AI that wouldn’t let her finish her sentences, hundreds of thousands of recordings had already been made.

The Real Question: How Much Data Is Enough?
I want to be clear: I am not arguing that AI should not be used in schools. Quite the opposite. Personalized learning has genuine potential to help teachers understand their students better and to give every child support that matches their specific needs.
A well-designed AI reading tutor can identify a child’s struggles with phonics long before a human teacher, managing 28 other students, ever could. It can track progress over months, spot patterns invisible to the human eye, and free up teachers to do what only humans can do: connect, inspire, and mentor.
But the question we need to ask is not “Can AI personalize learning?” It is this: How much of a child’s personal data should we give AI in order to personalize it?
Does a reading tutor need to record and retain a child’s voice to function? Or could it process the audio locally, extract the relevant metrics (fluency, accuracy, comprehension), and discard the raw recording immediately? These are design choices, not technical inevitabilities.
The difference between these two approaches is the difference between personalized learning and permanent surveillance. One serves the child. The other serves the data economy.
What Needs to Change: 5 Principles for Ethical AI in Education
If we are going to deploy AI tools in classrooms, and I believe we should, carefully, then we need a new set of principles that protect children before the technology scales further:
- Transparency Before Deployment. Schools must inform parents what data an AI tool collects, why it is collected, where it goes, and how long it is kept, before the tool is introduced to children, not after. The informed consent process cannot be a buried checkbox in a back-to-school packet.
- Data Minimization by Default. AI tools should collect only the data they genuinely need to function. If a reading assessment can work by processing audio locally and discarding the recording, that should be the only mode available. “Collect everything, decide later” has no place in schools.
- No AI Training on Student Data. Student data, especially biometric data like voice recordings, should never be used to train machine learning models without explicit, separate, opt-in parental consent. Educational data should not be raw material for the next generation of commercial AI products.
- Portability and Deletion Rights. When a child leaves a school or a district stops using a tool, all of that child’s data should be deleted, verifiably. Parents should have the right to request and confirm deletion at any time, without jumping through bureaucratic hoops.
- Independent Audits. Edtech companies making claims about data privacy should be subject to independent, third-party audits, just as financial institutions are. A company’s own privacy policy is not enough. Their claims must be verified.
Technology companies may see data as an asset. Schools should see children’s data as a responsibility.
Final Thoughts
The Amira story in New Mexico is bigger than one reading tool. As I reflected after the Salamanca AI teaching assistant experiment, when we introduce AI into classrooms, we are not just adding a tool, we are reshaping the relationship between students, teachers, and institutions. The Amira case is a preview of a future that is arriving faster than our laws, our policies, and our public understanding can keep up with.
AI in education is not going away. Nor should it. The promise of tools that help every child learn at their own pace, with support tailored to their specific needs, is genuinely exciting. But that promise must not be built on the invisible exploitation of children’s data, a concern I raised when I asked whether AI is really improving education or just adding a layer of technological complexity to existing problems.
The conversation that education leaders, parents, teachers, and technology companies need to have is not whether AI should be in classrooms. That ship has sailed. The conversation we need is about the terms of the relationship.
The Christian Science Monitor’s investigation into education’s AI “gold rush” captures the urgency of this moment. When everyone rushes to stake a claim, safeguards are often written after the damage is done. In this rush, the gold is not underground. It is the personal data of millions of children, and the people collecting it may not be the ones who bear the consequences.
Who owns the data? Who controls it? Who benefits from it? And who is left holding the consequences when something goes wrong?
If the answer to any of those questions is “we’ll figure it out later,” then we are not doing personalized learning. We are running an uncontrolled experiment, and our children are the subjects.
