No, not without permission or a copyright exception. Department for Education guidance says schools and colleges must not allow or cause students' original work to be used to train generative AI models. Permission comes from the copyright owner: the student, or a parent or guardian where the student is a child.
What does Department for Education guidance say about training on student work?
It sets a prohibition with a narrow exit. The Department for Education's guidance on generative AI in education tells schools and colleges they "must not allow or cause students' original work to be used to train" generative AI models, unless they have permission or an exception to copyright applies.
Two parts of that sentence do the work. "Allow or cause" covers passive failure as much as an active decision. If a supplier's default setting feeds uploaded coursework into a training set and nobody in the setting turned it off, the setting allowed it. And "permission or an exception" means there are exactly two lawful routes, and you have to be able to name which one you are relying on.
The same guidance tells settings to take their own legal advice on intellectual property. I would follow that. What follows is how to structure the question before you spend money on it, not a substitute for a lawyer who knows your contracts.
Who owns the copyright in a student's work?
The student does, in almost every case. Under section 11 of the Copyright, Designs and Patents Act 1988, the author of a work is the first owner of copyright in it. Age is not a qualifying condition. A nine year old who writes a story owns the copyright in that story from the moment it is written down.
The institution does not acquire ownership by setting the homework, marking it, displaying it on a wall or storing it on the network. Possession of a file is not a licence to use it.
That determines who you ask. For a student over 18, you ask the student. For a child, permission has to come from someone who can give it on the child's behalf, in practice a parent or guardian, which reflects the general difficulty of binding a minor to an agreement. Former students still own their work, so a leaver's portfolio sitting on a shared drive from three years ago is not free material. Neither is an archive of marked scripts.
What counts as original student work?
More than essays. The originality threshold in UK copyright is low: the work has to be the student's own intellectual creation rather than a copy, and it has to be recorded in some form. Coursework, homework, creative writing, artwork, design drawings, photographs, source code, recorded music and video, a dissertation, a portfolio. All of it is capable of protection. Ideas and bare facts are not protected. The expression of them is.
Staff material sits differently, and this is where institutions get the answer wrong in both directions. Where a teacher creates a lesson plan or a resource in the course of their employment, the employer is generally the first owner under the same Act, subject to what the employment contract actually says. So a trust may be able to consent to its own staff material being used for training when it cannot consent for pupil work. Read the contracts before assuming either way.
When does an exception to copyright apply?
Rarely, for training. The gov.uk list of exceptions to copyright is specific, and each exception is bounded by purpose. Two get raised in procurement conversations.
The first is illustration for instruction, which permits limited copying for teaching. It is about using work in front of a class, not about assembling a dataset. The second is the text and data mining exception, which allows copies for computational analysis where the purpose is non-commercial research and the person making them already has lawful access. A supplier improving a commercial product is not conducting non-commercial research.
If somebody tells you an exception covers it, ask which exception, in which section, and how the purpose test is met. Government policy on copyright and AI training is still contested and under review, so the position may move. Today, the safe working assumption for a school or college is that no exception permits training a commercial model on pupil work.
How do we check what a supplier does with uploaded work?
Read the contract, not the marketing page. You are looking for a written term saying that customer content is not used to train, retrain or fine-tune any model, and is not used to improve the service generally. The absence of a promise is not a promise.
Then get five answers in writing and keep them:
- Is training on our content off by default, or off only if we ask?
- Does that bind sub-processors and any model provider sitting behind the product?
- Where is our content processed and stored, and for how long?
- Can any person read our content, and under what circumstances?
- What happens to content already uploaded if we change the setting or leave?
Insist on contractual wording rather than a help centre article. A support page can be rewritten overnight. A term cannot. Where pupil personal data is in play, the ICO's guidance on AI and data protection covers a separate set of duties on lawful basis, transparency and impact assessment. Those run alongside the copyright question. Neither answer satisfies the other.
What should a permission request look like if we need one?
Specific, limited and genuinely refusable. A blanket line in a home-school agreement allowing the setting to use pupil work "for any purpose, including technology development" fails the basic test, because nobody signing it understands what they agreed to.
State whose work it is, which pieces, what will be done with it, whose model is being trained, whether the work leaves the institution, whether it can be withdrawn later and what withdrawal actually achieves, and how long the permission lasts. Say plainly that refusing affects nothing about marks, references or a place on a course. Record each answer against the individual, not as a tick on a class list.
And be honest about the hard part. If a student withdraws permission after training has happened, removing their contribution from a trained model is not a simple deletion. That difficulty is itself an argument for not needing the permission in the first place.
Does an AI that answers from documents train on them?
No, and this distinction matters more than anything else here. Retrieval is not training. A system can search your documents, quote them and cite them at the moment a question is asked, without any of that content altering a model's weights. Remove a document from the index and it stops informing answers from that point. Nothing was absorbed.
That is how we built the Mickai Sovereign Intelligence Operating System. Private knowledge bases, which we call brains, sit on hardware the institution owns, offline capable, with no data egress. The model reads what it is pointed at, and the reading leaves a record. The Open Audit Record seals every consequential action under ML-DSA-65, the post-quantum signature scheme NIST published as FIPS 204 in 2024, and an auditor can verify an exported record offline with a public key, using tools that are not ours. That is tamper-evident, which is a weaker and more useful claim than the one vendors usually reach for. We cannot stop someone altering a record. We can make sure an altered record fails verification, and that is what an auditor needs. Consequential actions wait for a named person to approve them.
One honest limit. A local OCR runtime has read scanned PDFs in controlled tests, and extraction and ingestion integration into SIOS is still being completed, so if your coursework archive is paper, ask about that before planning around it.
None of this is an argument against cloud. For work carrying no regulatory weight, cloud remains the sensible tool. The argument is against the assumption that a regulated organisation must rent its intelligence, ship its records offsite and take a vendor's word for what happened to them.
What should the institution's AI policy say about this?
One paragraph a head of department can act on without calling anybody. Name the default: student work is not used to train any AI model. Name who can vary that, and make it a person rather than a committee. Require that every new tool is checked against the training question before it is used, not after. Keep a register of approved tools with the relevant contract clause referenced beside each one. Say what staff should do about a tool a student brought in themselves, because that is the common case.
Then deal with what has already happened. Establish which tools have had coursework uploaded to them, what their terms said at the time, and write the answer down. There is a second direction to this as well: secondary infringement, where outputs from a model trained on unlicensed material are then used in the setting. So the check runs in both directions: what leaves, and what comes back.
Frequently asked questions
Can a school let an AI tool train on pupils' essays?
Not without permission or a copyright exception. Department for Education guidance tells schools and colleges not to allow or cause students' original work to be used to train generative AI models. Permission must come from the copyright owner, which is the pupil, or a parent or guardian for a child. A supplier's default setting counts as allowing it.
Who owns the copyright in a student's coursework?
The student. Under section 11 of the Copyright, Designs and Patents Act 1988 the author of a work is its first owner, and there is no minimum age. Setting the task, marking the work or storing it on the school network transfers nothing. Former students keep copyright in work they produced while enrolled, including portfolios left behind.
Does uploading work to an AI tool count as training it?
Not necessarily. A system can retrieve and quote a document at the moment a question is asked without that content altering any model weights. Whether your upload trains anything depends entirely on the supplier's terms and settings. Ask for a contractual term stating that customer content is never used to train, retrain or fine-tune a model.
Do teachers' lesson plans have the same protection?
They carry copyright, but ownership usually differs. Where a lesson plan is created by an employee in the course of employment, the employer is generally first owner under the same Act, subject to the employment contract. So a school or trust may be able to consent for staff material when it cannot consent for pupil work. Check contracts first.
How do we check a supplier's terms on training?
Read the contract, not the marketing page. Look for a term stating that customer content is not used to train, retrain, fine-tune or otherwise improve any model, and that the term binds sub-processors and any underlying model provider. Then ask in writing whether training is off by default, where content is stored, and what happens when you leave.
Related briefings
Education
- Private AI for University Staff: Keeping Data In House
- AI Governance for Further Education Colleges: Checklist
Data protection and UK GDPR
- UK GDPR and AI: Does Your Data Have to Stay in the UK?
- UK Data Storage vs AI Processing: What Is the Difference
Part of a series of 60 briefings on deploying and governing AI in UK regulated organisations, archived with a DOI at 10.5281/zenodo.22975756.
Evaluating AI for a regulated organisation? Mickai runs on hardware you own, offline. Consequential actions wait for a named person to approve them, and what the AI did is sealed into a signed record an auditor can check without us. Applications for the invitation-only closed beta are open. Apply for the closed beta.
Written by Micky Irons, founder and chief executive of Mickai LTD.
Top comments (0)