1 of 27

GenAI Output and Copyright�

LLMxLaw Conference: Trusting AI in High-Stakes Legal Domain

Judge Business School, June 2026

Henning Grosse Ruse - Khan

2 of 27

Outline

  1. The context: AI training and genAI’s capacity to ‘create’
  2. © protection for genAI outputs
  3. © liability for genAI output
  4.  genAI and the future of (human?) creativity

3 of 27

The context: AI training and genAI’s capacity to ‘create’

4 of 27

The importance of data for AI training

‘The race to lead A.I. has become a desperate hunt for the digital data needed to advance the technology.’ To obtain that data, tech companies including OpenAI, Google and Meta have cut corners, ignored corporate policies and debated bending the law, according to an examination by The New York Times. (….)

[N]ews stories, fictional works, message board posts, Wikipedia articles, computer programs, photos, podcasts and movie clips — has increasingly become the lifeblood of the booming A.I. industry.’

(NY Times, 6 April 2024)

5 of 27

Training AI on © Infringing Content...

Kadrey vs Meta, 18 December 2024 Filing

The Guardian, 10 Jan 2025

6 of 27

Use of © material for training AI

© issues arising:

  • Some data used for AI training can be sourced via licensing
  • As © material is everywhere online, relying on publicly available sources on the internet means that AI developers will use large amounts of © content without express permission
  • Since such use likely constitutes copying, and since there are no strong arguments for an implied license to use © material online for training AI, AI developers need to rely on © exceptions & limitations
  • Since data is likely to be collected from sources around the world, questions arise whose © law governs

7 of 27

8 of 27

EU: TDM exceptions in the DSM Directive

Art.4 - Exception or limitation for text and data mining

1.   Member States shall provide for an exception (…) for reproductions and extractions of lawfully accessible works and other subject matter for the purposes of text and data mining.

2.   Reproductions and extractions made pursuant to paragraph 1 may be retained for as long as is necessary for the purposes of text and data mining.

3.   The exception or limitation provided for in paragraph 1 shall apply on condition that the use of works and other subject matter referred to in that paragraph has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online.

9 of 27

UK: Copyright Politics

10 of 27

© protection for genAI outputs

11 of 27

© for an AI-generated image: Li vs Liu

Intellectual achievement” refers to the result of a human being’s intellectual activities. Plaintiff’s intellectual activities are evinced from the conception to the final creation of the disputed image. Through Stable Diffusion, Plaintiff selected over 150 prompts, arranged their order and set specific parameters. He continued to adjust and modify those prompts and parameters until the final image aligned with his conception. These steps sufficiently demonstrate that the disputed image was created as a result of Plaintiff’s intellectual inputs.

Furthermore, “originality” is manifested in Plaintiff’s personalized choices and aesthetic judgement throughout the generation process. This involved not only the selection and arrangement of the prompts and parameters, but also the refinement of the final output. Therefore, such “intellectual achievements” transcended the mere “mechanical” ones that are devoid of originality. (see Kluwer © Blog)

12 of 27

© for this comic, made with (or by?) AI?

12

13 of 27

US Copyright Office decision, 02/2023

Based on the record before it, the Office concludes that the images generated by Midjourney contained within the Work are not original works of authorship protected by copyright. Though she claims to have guided” the structure and content of each image, the process described (…) makes clear that it was Midjourney—not Kashtanova—that originated the “traditional elements of authorship” in the images.

First, she entered a text prompt to Midjourney, which she describes as “the core creative input” for the image. Next, “Kashtanova then picked one or more of these output images to further develop.” She then “tweaked or changed the prompt as well as the other inputs provided to Midjourney” to generate new intermediate images, and ultimately the final image.

Rather than a tool that Ms. Kashtanova controlled and guided to reach her desired image, Midjourney generates images in an unpredictable way. Accordingly, Midjourney users are not the “authors” for copyright purposes of the images the technology generates.

14 of 27

V. Conclusion

Based on the fundamental principles of copyright, the current state of fast-evolving technology, and the information received in response to the NOI, the Copyright Office concludes that existing legal doctrines are adequate and appropriate to resolve questions of copyrightability. Copyright law has long adapted to new technology and can enable case-by-case determinations as to whether AI-generated outputs reflect sufficient human contribution to warrant copyright protection. As described above, in many circumstances these outputs will be copyrightable in whole or in part—where AI is used as a tool, and where a human has been able to determine the expressive elements they contain. Prompts alone, however, at this stage are unlikely to satisfy those requirements.

15 of 27

AG Munich, 02 2026

‘[C]opyright protection depends on whether the product "reflects the personality of its author by expressing their free creative decisions." (para. 18). (…)

the court stated that for copyright to subsist, the AI model must be "closer to a mere tool (Hilfsmittel) than to an independent instrument of creation." (para. 21). Also, it is not enough to merely "trigger" a process or select from multiple outputs.

In Levola Hengelo (C-310/17), the CJEU ruled that the work itself must be identifiable with "sufficient objectivity." Here, the Munich court takes that European requirement and applies it upstream to the prompt. The court’s logic is: If the prompt is so vague ("make it artistic") that the AI has a billion ways to interpret it, the human has not "objectively defined" the output. The output is then a product of the machine's randomness, not the human's identifiable intent.

Protection only arises if the creative elements within the prompt "dominate the output so much that the object can be seen as the author's own original creation." (para. 19-21). The court held that "merely generally formulated, open-ended instructions" leave the actual "design decision" to the AI, thus failing the test for authorship. (Duhanic)

16 of 27

UK provisions on computer-generated works: �an interesting model?

CDPA, s 9(3): In the case of a literary, dramatic, musical or artistic work which is computer-generated, the author shall be taken to be the person by whom the arrangements necessary for the creation of the work are undertaken.

CDPA, s 178: “computer-generated”, in relation to a work, means that the work is generated by computer in circumstances such that there is no human author of the work

CDPA, s 12(7): If the work is computer-generated the above provisions do not apply and copyright expires at the end of the period of 50 years from the end of the calendar year in which the work was made.

CDPA, s.78(2)(c), s. 81(2): no ‘moral’ rights

17 of 27

Is the UK Provision A Useful Model for Protection of AI Generated Works?

Not really.

  1. Fails to Solve Originality Problem. Report on Copyright and Artificial Intelligence (18 March 2026): ‘This contradiction has led some to question whether the provision could ever apply in practice. In our view, it is unlikely that a court would conclude that it can never apply, as Parliament clearly intended the provision to have an effect. But it is unclear in the absence of case law how an “original” yet wholly machine-authored work would be defined.’
  2. Has Not Produced Clarity: Who ‘makes the arrangements’?

🡪 ‘[I]n the absence of evidence of its ongoing value, we propose that [sec.9.3 CDPA] should be removed’

(However, protection of genAI output under related rights – as sound recording (sec.5A) or film (sec.5B) – remains an option)

18 of 27

© liability for genAI output

19 of 27

Memorisation as indicator for © breach?

20 of 27

Who is responsible? User or AI provider?

🡪 AI providers might be contributing to infringements by users, or infringe directly

21 of 27

… As showing AI training to be © infringing faces hurdles of applicable law and potential © exceptions (fair use, TDM), some right holders are shifting attention on AI models, and the output they generate

Direct infringement: Does the trained genAI Model include copies?

22 of 27

Does the trained (gen)AI Model include copies?

Stable Diffusion does not itself store the data on which it was trained … Rather than storing their training data, diffusion models learn the statistics of patterns… It is impossible to store all training images in the weights… LAION-5B ~220TB vs 3.44GB model weights.” [552–554]

"An infringing copy must be a copy… I cannot see how an article can be an infringing copy if it has never consisted of/stored/contained a copy. In Sony v Ball the RAM chip was only an infringing copy while it contained the copy” [584, 587]

🡪 While memorisation can occur, Getty did not allege that any © work was memorised or stored in the weights:

"There is no evidence of any Copyright Work having been ‘memorized’… and no evidence of any image having been derived from a Copyright Work.” [559, see also 560]

🡪 Even though training elsewhere involves reproductions, weights that never contain a copy are not an ‘infringing copy’:

"Is an AI model which derives from a training process itself an infringing copy? In my judgment, it is not… by the end of that process the model does not store any of those works… The model weights… have never contained or stored an infringing copy."�[items 599–600]

23 of 27

Does the trained (gen)AI Model include copies?

In the opinion of the Chamber, the song lyrics in dispute are reproducible in the defendant‘s language models 4 and 4o. It is known from information technology research that training data can be contained in language models and can be extracted as outputs. This is referred to as memorisation. This occurs when the language models not only extract information from the training data set during training, but also completely adopt the training data in the parameters specified after training. Such memorisation was determined by comparing the song lyrics contained in the training data with the reproductions in the outputs. Given the complexity and length of the song lyrics, coincidence can be ruled out as the cause of the reproduction of the song lyrics.

Memorisation means that embodiment, as a prerequisite for the copyright reproduction of the song lyrics in dispute, is given by data in the specified parameters of the model. The song lyrics in dispute are reproducibly defined in the models. According to Art. 2 of the InfoSoc Directive, reproduction ‘in any manner and in any form’ is considered to exist. The specification in mere probability values is irrelevant in this context. New technologies such as language models would be covered by the reproduction right under Art. 2 of the InfoSoc Directive and Section 16 of the UrhG. According to the case law of the Court of Justice of the European Union, indirect perceptibility is sufficient for reproduction, which is given if the work can be perceived using technical aids.

24 of 27

Contributory infringement: towards intermediary liability of AI providers?

  • Authorisation liability, sec.16.2 CDPA: Does providing a tool which allows users to create potentially © infringing content amount to authorising the user to infringe? 🡪 See HL in CBS v Amstrad, 1988
  • Secondary infringement in the UK: importing, possessing / dealing with, or providing means for making, an ‘infringing copy‘ (sec.22-24 CDPA) 🡪 See Getty v Stability AI, now at the CA...
  • Safe harbours (e.g. for host providers under Art.6 DSA) protecting (at least) online platforms that embedd AI tools in their services?
    • See recent CJEU GC judgment WebGroup / Coyote: 🡪 no host provider safe harbours if, beyond mere categorisation/indexation, an algorithm ‘determines, in the interest of the operator or its service, under what conditions, how and in which order of priority’ the hosted information is made available. [106-112]
    • Hence in cases of AI-based content moderation / recommender systems, platform ‘exercises control over that information’ and thus falls outside hosting safe harbour… 🡪 Implications for genAI providers?

25 of 27

genAI and the future of (human?) creativity

26 of 27

Increasing amounts of online content are AI-generated

  • What follows for the default assumption about © protection?
  • And how to ensure that creativity is not outsourced to AI?

27 of 27

Contamination of the online data environment by AI outputs?�Is access to uncontaminated (human) data an ‘essential resource’ and public policy concern that the law should address before it is too late?��If so, how might such access be afforded in a fair and equitable manner, while minimising the risks of harm from such access?