GenAI Output and Copyright�
LLMxLaw Conference: Trusting AI in High-Stakes Legal Domain
Judge Business School, June 2026
Henning Grosse Ruse - Khan
Outline
The context: AI training and genAI’s capacity to ‘create’
The importance of data for AI training
‘The race to lead A.I. has become a desperate hunt for the digital data needed to advance the technology.’ To obtain that data, tech companies including OpenAI, Google and Meta have cut corners, ignored corporate policies and debated bending the law, according to an examination by The New York Times. (….)
[N]ews stories, fictional works, message board posts, Wikipedia articles, computer programs, photos, podcasts and movie clips — has increasingly become the lifeblood of the booming A.I. industry.’
(NY Times, 6 April 2024)
Training AI on © Infringing Content...
Kadrey vs Meta, 18 December 2024 Filing
The Guardian, 10 Jan 2025
Use of © material for training AI
© issues arising:
EU: TDM exceptions in the DSM Directive
Art.4 - Exception or limitation for text and data mining
1. Member States shall provide for an exception (…) for reproductions and extractions of lawfully accessible works and other subject matter for the purposes of text and data mining.
2. Reproductions and extractions made pursuant to paragraph 1 may be retained for as long as is necessary for the purposes of text and data mining.
3. The exception or limitation provided for in paragraph 1 shall apply on condition that the use of works and other subject matter referred to in that paragraph has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online.
UK: Copyright Politics
© protection for genAI outputs
© for an AI-generated image: Li vs Liu
“Intellectual achievement” refers to the result of a human being’s intellectual activities. Plaintiff’s intellectual activities are evinced from the conception to the final creation of the disputed image. Through Stable Diffusion, Plaintiff selected over 150 prompts, arranged their order and set specific parameters. He continued to adjust and modify those prompts and parameters until the final image aligned with his conception. These steps sufficiently demonstrate that the disputed image was created as a result of Plaintiff’s intellectual inputs.
Furthermore, “originality” is manifested in Plaintiff’s personalized choices and aesthetic judgement throughout the generation process. This involved not only the selection and arrangement of the prompts and parameters, but also the refinement of the final output. Therefore, such “intellectual achievements” transcended the mere “mechanical” ones that are devoid of originality. (see Kluwer © Blog)
© for this comic, made with (or by?) AI?
12
US Copyright Office decision, 02/2023
Based on the record before it, the Office concludes that the images generated by Midjourney contained within the Work are not original works of authorship protected by copyright. Though she claims to have “guided” the structure and content of each image, the process described (…) makes clear that it was Midjourney—not Kashtanova—that originated the “traditional elements of authorship” in the images.
First, she entered a text prompt to Midjourney, which she describes as “the core creative input” for the image. Next, “Kashtanova then picked one or more of these output images to further develop.” She then “tweaked or changed the prompt as well as the other inputs provided to Midjourney” to generate new intermediate images, and ultimately the final image.
Rather than a tool that Ms. Kashtanova controlled and guided to reach her desired image, Midjourney generates images in an unpredictable way. Accordingly, Midjourney users are not the “authors” for copyright purposes of the images the technology generates.
V. Conclusion
Based on the fundamental principles of copyright, the current state of fast-evolving technology, and the information received in response to the NOI, the Copyright Office concludes that existing legal doctrines are adequate and appropriate to resolve questions of copyrightability. Copyright law has long adapted to new technology and can enable case-by-case determinations as to whether AI-generated outputs reflect sufficient human contribution to warrant copyright protection. As described above, in many circumstances these outputs will be copyrightable in whole or in part—where AI is used as a tool, and where a human has been able to determine the expressive elements they contain. Prompts alone, however, at this stage are unlikely to satisfy those requirements.
AG Munich, 02 2026
‘[C]opyright protection depends on whether the product "reflects the personality of its author by expressing their free creative decisions." (para. 18). (…)
the court stated that for copyright to subsist, the AI model must be "closer to a mere tool (Hilfsmittel) than to an independent instrument of creation." (para. 21). Also, it is not enough to merely "trigger" a process or select from multiple outputs.
In Levola Hengelo (C-310/17), the CJEU ruled that the work itself must be identifiable with "sufficient objectivity." Here, the Munich court takes that European requirement and applies it upstream to the prompt. The court’s logic is: If the prompt is so vague ("make it artistic") that the AI has a billion ways to interpret it, the human has not "objectively defined" the output. The output is then a product of the machine's randomness, not the human's identifiable intent.
Protection only arises if the creative elements within the prompt "dominate the output so much that the object can be seen as the author's own original creation." (para. 19-21). The court held that "merely generally formulated, open-ended instructions" leave the actual "design decision" to the AI, thus failing the test for authorship. (Duhanic)
UK provisions on computer-generated works: �an interesting model?
CDPA, s 9(3): In the case of a literary, dramatic, musical or artistic work which is computer-generated, the author shall be taken to be the person by whom the arrangements necessary for the creation of the work are undertaken.
CDPA, s 178: “computer-generated”, in relation to a work, means that the work is generated by computer in circumstances such that there is no human author of the work”
CDPA, s 12(7): If the work is computer-generated the above provisions do not apply and copyright expires at the end of the period of 50 years from the end of the calendar year in which the work was made.
CDPA, s.78(2)(c), s. 81(2): no ‘moral’ rights
Is the UK Provision A Useful Model for Protection of AI Generated Works?
Not really.
🡪 ‘[I]n the absence of evidence of its ongoing value, we propose that [sec.9.3 CDPA] should be removed’
(However, protection of genAI output under related rights – as sound recording (sec.5A) or film (sec.5B) – remains an option)
© liability for genAI output
Memorisation as indicator for © breach?
Who is responsible? User or AI provider?
🡪 AI providers might be contributing to infringements by users, or infringe directly
… As showing AI training to be © infringing faces hurdles of applicable law and potential © exceptions (fair use, TDM), some right holders are shifting attention on AI models, and the output they generate…
Direct infringement: Does the trained genAI Model include copies?
Does the trained (gen)AI Model include copies?
Stable Diffusion does not itself store the data on which it was trained … Rather than storing their training data, diffusion models learn the statistics of patterns… It is impossible to store all training images in the weights… LAION-5B ~220TB vs 3.44GB model weights.” [552–554]
"An infringing copy must be a copy… I cannot see how an article can be an infringing copy if it has never consisted of/stored/contained a copy. In Sony v Ball the RAM chip was only an infringing copy while it contained the copy” [584, 587]
🡪 While memorisation can occur, Getty did not allege that any © work was memorised or stored in the weights:
"There is no evidence of any Copyright Work having been ‘memorized’… and no evidence of any image having been derived from a Copyright Work.” [559, see also 560]
🡪 Even though training elsewhere involves reproductions, weights that never contain a copy are not an ‘infringing copy’:
"Is an AI model which derives from a training process itself an infringing copy? In my judgment, it is not… by the end of that process the model does not store any of those works… The model weights… have never contained or stored an infringing copy."�[items 599–600]
Does the trained (gen)AI Model include copies?
In the opinion of the Chamber, the song lyrics in dispute are reproducible in the defendant‘s language models 4 and 4o. It is known from information technology research that training data can be contained in language models and can be extracted as outputs. This is referred to as memorisation. This occurs when the language models not only extract information from the training data set during training, but also completely adopt the training data in the parameters specified after training. Such memorisation was determined by comparing the song lyrics contained in the training data with the reproductions in the outputs. Given the complexity and length of the song lyrics, coincidence can be ruled out as the cause of the reproduction of the song lyrics.
Memorisation means that embodiment, as a prerequisite for the copyright reproduction of the song lyrics in dispute, is given by data in the specified parameters of the model. The song lyrics in dispute are reproducibly defined in the models. According to Art. 2 of the InfoSoc Directive, reproduction ‘in any manner and in any form’ is considered to exist. The specification in mere probability values is irrelevant in this context. New technologies such as language models would be covered by the reproduction right under Art. 2 of the InfoSoc Directive and Section 16 of the UrhG. According to the case law of the Court of Justice of the European Union, indirect perceptibility is sufficient for reproduction, which is given if the work can be perceived using technical aids.
Contributory infringement: towards intermediary liability of AI providers?
genAI and the future of (human?) creativity
Increasing amounts of online content are AI-generated
Contamination of the online data environment by AI outputs?��Is access to uncontaminated (human) data an ‘essential resource’ and public policy concern that the law should address before it is too late?��If so, how might such access be afforded in a fair and equitable manner, while minimising the risks of harm from such access?