DEV Community

Cover image for Protecting the IP in a Generative World
Andrew Stevens for Sakura Sky

Posted on Originally published at sakurasky.com

Protecting the IP in a Generative World

A studio found its own back catalogue inside the training data of a public model. Someone had run a handful of its older titles through a detection tool, and the tool reported, with reasonable confidence, that the model had seen them. The legal team moved quickly, because a decades-old library is the company's principal asset and defending it is the job. The engineering team was asked a simpler question and could not answer it. From the studio's own systems, could it prove when each of those works was created, what it was derived from, and whether anyone had ever been licensed to ingest it. The plain answer was no. The studio owned the content and could not produce the provenance.

That gap is the subject of this post. The studio's real exposure went well beyond a model having trained on its work. Nothing in its own architecture could establish the origin, ownership, and rights history of the things it makes, on demand and in a form an outside party would accept. The wider series argued that trust is something a system produces rather than asserts (see Trust Is an Engineering Output). For a media business, provenance is the specific shape that takes, and generative models are the forcing function that has finally made its absence expensive. This post works backwards from the discovery: what the old protections were built to do, what protection now demands, and the layer that closes the gap.

The back-catalogue discovery

Start with what the studio did have, because it was not nothing. Its digital rights management was intact. Streams were encrypted, playback was licensed, access to the finished assets was controlled, and outbound video carried watermarks. By the standard of the threat that DRM was built for, the studio was well defended. None of it touched the problem in the room.

What the studio lacked was a record, attached to each asset and maintained by its systems, of when the work was made, who contributed to it, what source material it drew on, and what rights sat over it. That information existed, scattered across contracts in a document store, production notes, and the memories of people who had moved on. It did not exist as data the studio could query. So when the question became "prove this is yours and prove what happened to it," the studio was reduced to reconstructing its own history by hand, which is exactly the position the earlier posts in this programme described as evidence that was never engineered.

This is the shape most media businesses are in, and it is worth being precise about why it is uncomfortable. Without that record, the studio could not prove ingestion had happened, could not cheaply assert ownership, and could not act on the opt-out mechanisms the law now provides, because it had no machine-readable way to express its rights across a library built over decades. The content was owned. The provenance was not producible. Those are different properties, and only the second one was being tested.

What DRM used to mean

Work backwards to how the protections were built, because they were built well for a threat that has changed. Digital rights management was designed to stop one thing: the unauthorised copying and redistribution of a finished asset. Encryption at rest and in transit, licence servers that decide who may play what, access controls, and watermarks that trace a leak back to its source. The protected object was the copy, and the moment of risk was consumption. For piracy, this was the right design, and for piracy it still largely works.

It was silent, though, on the things that now matter most. Encryption protects the delivery path, so it does block a crawler from ingesting that path, but the training-data problem mostly lives elsewhere: in the trailers, clips, and released titles that circulate in the clear, and in the studio's ability to express its rights and prove its own lineage. DRM protects the outbound stream. It says nothing about how a work may be used for training, nor about the origin story of the work behind the stream.

This is why DRM modernisation is not a matter of stronger encryption or tighter licence enforcement. The object that now needs protecting is not only the outbound stream. It is the provenance of the work itself, the verifiable account of what it is, when it was made, and what may lawfully be done with it. That account is data, and it has to be engineered as data.

What model-era IP protection requires

Once the threat moves from copying to ingestion, content IP protection needs three capabilities that traditional media rights management never had to provide, and the law on both sides of the Atlantic is moving toward them, even as the central US question, whether training counts as fair use, stays unresolved in the courts.

The first is machine-readable rights reservation. European copyright law gives rightsholders a text-and-data-mining opt-out, but it only functions if the reservation is expressed in a machine-readable form that a crawler can detect (European Parliament and Council, 2019). The AI Act then obliges providers of general-purpose models to put in place a policy to comply with that reservation and to publish a sufficiently detailed summary of the content used to train the model (European Parliament and Council, 2024). An opt-out a studio cannot express at scale, across a whole catalogue, is a right it cannot actually exercise.

The second is provable ownership. The United States arrives at a related point from the other direction: in a pre-publication report, the US Copyright Office takes the view that assembling a training dataset from copyrighted works implicates the reproduction right, because it involves downloading, storing, and copying those works (U.S. Copyright Office, 2025). Registration and a clean chain of title remain the legal instruments that let a studio sue and recover, so the provenance layer does not replace them. It feeds them, giving the studio a queryable evidence base for what it owned and when, which is the thing a claim or a defence eventually turns on.

The third is content provenance. The industry has converged on a standard for this, the C2PA content credentials specification, which attaches tamper-evident provenance and edit history to a media asset so its origin travels with it rather than living in a separate database (C2PA, 2026). Generative AI copyright disputes are often, underneath the legal language, arguments about who can prove what about a file, and content credentials are an attempt to make that provable at the level of the asset.

This is not hypothetical, and music is where it has been tested first. In July 2026 the Munich Regional Court found the AI music service Suno liable for copyright infringement in a case brought by the German collecting society GEMA, holding that specific protected works remained reproducible inside the models and that the outputs reproduced them, and that training carried out in the United States did not place it beyond German law (Music Business Worldwide, 2026). The decision is first-instance and under appeal, but it turns on exactly the question a provenance layer exists to answer: can a rightsholder show that a particular work was used, and that a particular output reproduced it. The studios watching that case are asking whether they could prove the same thing about their own catalogues, and most cannot.

The provenance layer

Move forward from the diagnosis and the resolution is an IP-aware data layer, engineered before the dispute rather than assembled during it.

It does three things, and they are the implementations of the capabilities above. First, it makes provenance a property of every asset: content credentials are attached at the point of creation, and the studio keeps an internal record of each work's creation date, contributors, source and derivation, and attached rights. That record is made tamper-evident through cryptographic signing and hash-linking, so a later alteration is detectable, which is what turns a database into evidence. Second, it expresses machine-readable rights reservation consistently across the whole library, so a model training opt-out is something the studio asserts at scale rather than in principle. Third, it adds monitoring that checks whether the studio's works surface in public datasets or model outputs. That monitoring is probabilistic, the same reasonable-confidence signal the studio started with, so it works as an early-warning system rather than as proof, and describing it that way keeps the claim accurate. Building that layer out of a media company's scattered production and rights data is IP architecture, and it is a large part of what Sakura's Data & AI practice does for content businesses.

This is where the argument connects back to the rest of the programme. A provenance layer is the media form of evidence as an engineered property: the studio can answer, for any asset, where it came from and what may be done with it, on demand and with the signed record attached. That is data provenance doing the same job in a studio that a hash-linked transaction chain does in a bank, and it turns content provenance from a compliance aspiration into a running feature of the platform.

What happens when this is not engineered

The cost of skipping this is not only the litigation the studio is now in. It is broader, and some of it is opportunity rather than risk.

Without the layer, a media company cannot readily prove ingestion happened, cannot enforce an opt-out it has no way to express, and cannot assert ownership without a manual reconstruction every time. Each of those is a live exposure as the use of copyrighted works for AI training data becomes contested. The part most businesses miss is on the other side of the ledger. A market for licensing catalogues to model builders is forming: in the same music-industry fight, Warner Music settled its US infringement suit against Suno in late 2025 and signed a licensing partnership, even as Universal and Sony kept litigating (Music Business Worldwide, 2026). Those deals still close on corporate ownership and contractual warranties rather than on asset-level provenance alone, but a content owner that can show clean provenance diligences faster, carries less risk, and negotiates from a stronger position. The same missing layer that leaves the company exposed also leaves value on the table in the market its own content is helping to create.

The evidence that lets a media business defend its rights or license its catalogue is engineering output produced before the dispute, not a legal artefact produced after it. Building the record that lets a studio prove what it owns, express how it may be used, and show when a boundary was crossed is a media security and evidence problem before it is a legal one, and it is the ground Sakura's Security practice works on with a media company's data and rights teams.

References

Coalition for Content Provenance and Authenticity (C2PA), 2026. C2PA Technical Specification, Version 2.4. Available at: https://spec.c2pa.org/ [Accessed 18 August 2026].

Music Business Worldwide, 2026. Suno infringed copyright in GEMA case, German court rules. Reporting the Munich Regional Court first-instance judgment in GEMA v Suno, Case 42 O 763/25, 31 July 2026 (under appeal). Available at: https://www.musicbusinessworldwide.com/suno-infringed-copyright-in-gema-case-german-court-rules/ [Accessed 18 August 2026].

European Parliament and Council, 2019. Directive (EU) 2019/790 of the European Parliament and of the Council of 17 April 2019 on copyright and related rights in the Digital Single Market and amending Directives 96/9/EC and 2001/29/EC. Official Journal of the European Union, L 130, 17 May, pp. 92-125. Available at: https://eur-lex.europa.eu/eli/dir/2019/790/oj [Accessed 18 August 2026].

European Parliament and Council, 2024. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union, L 2024/1689, 12 July. Available at: https://eur-lex.europa.eu/eli/reg/2024/1689/oj [Accessed 18 August 2026].

U.S. Copyright Office, 2025. Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version). United States Copyright Office, Washington, DC, 9 May. Available at: https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf [Accessed 18 August 2026].

Top comments (0)