1 of 53

Copyright and AI

Reed C. Hepler

Hepler Consulting

Reed.hepler@gmail.com

heplerconsulting.com 

This presentation was made with the assistance of generative AI.

Ethics of AI Series

2 of 53

Learning Objectives

Identify and articulate key ethical concerns related to AI use in libraries.

Analyze how AI tools intersect with academic integrity and copyright.

Evaluate and apply best practices for AI use related to copyright.

Discuss implications of other ethical issues and actions through the lens of copyright law and related practices.

3 of 53

Perspectives on Generative AI in Research, Teaching, and Learning

Fear that use of GenAI primarily created new forms of unethical practices

Confidence that people inherently seek to use GenAI in effective and constructive ways

Fear that GenAI undermined systems and norms of information access and learning

Confidence that GenAI results in innovative products and workflows to enhance learning and research.

4 of 53

Three Cs of Generative AI

Copyright

who owns the rights?

Citation

which tool was used to create the material, and if necessary where did they get their information or model?

Circumspection

what hazards (moral, ethical, educational) should I manage?

5 of 53

Three Cs of Generative AI

Copyright

who owns the rights?

6 of 53

What are the Rights and Responsibilities of the Copyright Owner?

7 of 53

What are the Rights and Responsibilities of the User?

8 of 53

Which is Generative AI�Or... Is it Both or Neither?

A New Kind of User?

9 of 53

AI As a New Type of User of Copyrighted Materials

  1. Copyrighted material derivations have typically been direct user infringements.
  2. Generative AI tools are now mediating tools for any work that can be considered derivative.
  3. Did the AI tool infringe, or did the user infringe?
  4. Difference between training AI on copyrighted materials and using AI in a prompt

Data calculus-informed probabilitization vs. directed derivation

10 of 53

How Large Language Models Work

11 of 53

How Large Language Models Work

DATA

  • Conversion to tokens (“Strawberry Problem”)
  • Mathematical. Language, not logic.

MODEL

  • Pre-training (cultural bias)
  • Fine-tuning (training, human, political bias)

APPLICATIONS

  • Designed to seem/feel human
  • Artificial rapport / sycophant (reflecting our bias)
  • Everything is “hallucination” (not truth)

DANGERS

  • Propaganda and belief-shaping
  • Psychographic profiling

12 of 53

Two Conceptualizations of How LLMs Work

13 of 53

Two Conceptualizations of How LLMs Work

Stochastic Parrots/Octopi

    • Statistical
    • Probabilistic
    • Only repeats text, does not know meaning
    • stochastic: randomly determined; having a random probability distribution or pattern that may be analyzed statistically but may not be predicted precisely

World Domains

    • Making contextual connections on their own
    • Training, context, repetition, and user prompts reaffirm and alter existing connections
    • Creates “map” or “schema”

14 of 53

One View: Stochastic Parrots/Octopi

“Emily Bender is exactly right when she calls these models ‘stochastic parrots.’ No amount of increasing complexity can ever turn a nonrational, purely deterministic, mathematical process into a rational understanding of truth. Thus, ‘hallucinations.’ These models are something like a cultural mirror: if, when we gaze into them, what we see looks human, it's because we are human. It is decidedly NOT because the mirror has spontaneously become human.”

David W., comment on Lee, T., and Trott, S. (2024). “Large Language Models, explained with a minimum of math and jargon.” Understanding AI. Substack. https://www.understandingai.org/p/large-language-models-explained-with

15 of 53

One View: Stochastic Parrots/Octopi

“It turns out that if we provide enough data and computing power, language models end up learning a lot about how human language works simply by figuring out how to best predict the next word. The downside is that we wind up with systems whose inner workings we don’t fully understand.”

Lee, T., and Trott, S. (2024). “Large Language Models, explained with a minimum of math and jargon.” Understanding AI. Substack. https://www.understandingai.org/p/large-language-models-explained-with

16 of 53

Another View: World/Domain Models

David Chalmers, Prithviraj Ammanabrolu, and others propose a different model of thinking about AI tools.��They propose that rather than acting as stochastic parrots, AI models build data-supported connections between points, ideas, and keywords to create “domain-” or “world-models.”��This explains why, as AI tools are connected to the web and trained for longer periods of time, their errors are reduced.

“Sample domain model,” by Kishorekumar 62, is licensed under a CC BY SA 3.0 License.

17 of 53

Copyright

  • The owner
    • Creator
    • Custodian
    • other owner of these rights
  • Holds the title to
  • Intellectual property of a particular work.
    • Creation that was the result of the work of the mind of one or more people. 
      • Express ideas through any type of media

18 of 53

What Are the Rights in Copyright?

Exclusive rights to control 

    • duplications,
    • alterations, 
    • performance, 
    • display, and 
    • dissemination of a particular work, its expressions, and their manifestations
    • Includes derivations (unless it is a genuine parody)

19 of 53

What Can Be Copyrighted?

Literary works

Expressions of ideas through any type of media​

Musical works

Dramatic works​

Pictorial, graphic, and sculptural works

Motion pictures

Audiovisual works

Sound recordings​

Architectural works

Compilations and derivative works

20 of 53

What Cannot Be Copyrighted?

Ideas

Processes

Devices

Blank books, forms, charts, calendars, etc.

Laws and judicial opinions

Titles of works

Facts and data

Recipes

Works that have not been created by humans

Works of federal government employees

Public domain materials

21 of 53

Recommendations on Copyright and GenAI

Focus on existing Copyright Law

Mandate of Human Authorship

Does the Programmer Count as an Author?

Stay Informed

Prioritize Respecting IP

NEVER put copyrighted content in prompts.

Do Not Misrepresent AI Work as (Wholly) Human Work

AI is NOT a Workaround

22 of 53

Real-World Examples

23 of 53

24 of 53

What Arguments Would You Make For Or Against These Products?

25 of 53

Fair Use and New Arguments

26 of 53

Fair Use: Most Commonly-Used Argument

Factors to consider: 

How this affects use: 

The purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes 

Uses in nonprofit educational institutions are more likely to be fair use than works used for commercial purposes, but not all educational uses are fair use 

The nature of the copyrighted work. 

Reproducing a factual work is more likely to be fair use than a creative, artistic work such as a musical composition. Also, using an unpublished work would probably not be considered justifiable fair use. 

The amount and significance of the portion used in relation to the entire work 

Reproducing smaller portions of a work is more likely to be fair use than larger portions 

The effect of the use upon the potential market for or value of the copyrighted work 

Uses which have no or little market impact on the copyrighted work are more likely to be fair than those that interfere with potential markets 

27 of 53

Two Conceptual Models and Copyright

Stochastic Parrots

    • Statistical
    • Probabilistic
    • Copyright holders have little to worry about unless users maliciously prompt AI

World Domains

    • Making contextual connections on their own
    • Using user data to generate personalized things using transformed data
    • Not eligible for copyright protection, and
    • Also not infringing

28 of 53

Potential Fair Use Arguments for Training AI Tools

Non-Consumptive Use

    • Tool does not retain or transmit copyrighted material
    • Unless explicitly asked
    • Will only summarize or comment on the original work.
    • It uses copyrighted works to train on syntax and communication.
    • Stores metadata about what it “learns”

Non-Expressive Use

    • Making contextual connections on their own
    • Using user data to generate personalized things using transformed and contextualized data
    • Transforms preexisting ideas with new ideas provided by user
    • Not eligible for copyright protection, and
    • Also not infringing

29 of 53

Things to Keep in Mind: USCO Guidance and Misguided Practices

30 of 53

USCO Reports – Part 1

August 2024 – Part 1

    • Legislation is needed for deepfakes
    • Distinguishes between deepfakes and materials created to mimic another person’s creative style.
    • “Mimics” were held to be covered by preexisting laws.
    • Deepfakes required other legislation

31 of 53

USCO Reports – Parts 2 and 3

January – Part 2

    • Prompting an AI tool is not sufficient oversight
    • Altering, selecting, etc. is valid only for limited protection.
    • Case by case basis for works created through prompting (Kashtanova, Shupe)
    • “Arrangement, selection, curation, and other manipulation” of AI content can be copyrighted
    • Artists have to prove they did not use AI to make major decisions

May – Part 3

    • “GenAI tool training and use qualifies as fair use, BUT…
    • Majority of users are not using these tools that way… we need to establish rules.
    • “even paradigmatic fair uses… are often done for profit.”
    • in cases where there is a transformative purpose… copying of entire works may be reasonable.
    • “outputs are unlikely to substitute for expressive works used in training.”

32 of 53

Uploading Attachments is Not Training

  • One of the most misunderstood practice that users claim is “fair use”
  • Training involves transforming the data
  • Prompting with a document (or text grab!) deliberately and directly involves copyrighted information.
  • There is no fair use argument here, except for educators (and even that is tenuous)
  • Attachments are automatically viewed by the AI as something to derive from; phrases, sentences, and ideas are often lifted verbatim.

33 of 53

Malicious Prompts Lead to Copyright Violations

  • New York Times v. OpenAI
  • NYT used intentionally malicious prompts to coerce ChatGPT to generate articles verbatim, claimed copyrighted infringement.
  • Repeating copyrighted material is called “regurgitation” and is only done through extensive prompting.
  • Goes against OpenAI’s Terms of Use
  • This is the fault of the user.
  • NYT used an extreme case, but “malicious” does not always mean verbatim, half-article prompts.

34 of 53

AI-Copyright Trap

  • Paper by Carys Craig, professor at York University
  • “We should sidestep the copyright trap… in favor of more appropriate routes towards addressing the risks and harms of generative AI.
  • “Authorship is a fundamentally communicative act.”�“AI is categorically incapable of authoring original works of expression.”
  • “AI outputs will, themselves, have little economic value.”

35 of 53

Open Access and GenAI Training

  • Free to the public with no hidden costs or fees. 

  • Shared and duplicated

  • However, they cannot be revised, edited, clarified, or combined with other content.

  • Strictly open access licenses include Full Copyright licenses of materials that have been made publicly available, CC-BY-NC-ND license, CC-BY-ND license.

Creative commons license spectrum.svg was created by Shaddim and was licensed under a Creative Commons Attribution 4.0International license.

36 of 53

Open Access and GenAI Outputs

  • Humans nor AI can copyright AI outputs or products.
  • Are OA licenses a solution?
  • Creative Commons suggests using CC licenses as you would generally (preferably CC0 if there is little human involvement), some say to not use NC licenses.

Creative commons license spectrum.svg was created by Shaddim and was licensed under a Creative Commons Attribution 4.0International license.

37 of 53

Questions?

“In some cases, we learn more by looking for the answer to a question and not finding it than we do from learning the answer itself.” - Dallben, The Book of Three

38 of 53

Three Cs of Generative AI

Citation

which tool was used to create the tool, and if necessary where did they get their information or model?

39 of 53

Why Do We Cite?

40 of 53

Cite Your Sources!... and Tools

There is no set standard for citing AI tools. Even official suggestions by APA, MLA, and Chicago are just suggestions because of the constantly-changing perceptions of the nature of generative AI.

Know the AI use and citation policy for the school, class, and/or publication for which you are writing.

The ideal citation in any style should include:

  • Tool name and version (e.g., ChatGPT 3.5)
  • Time and date of usage
  • Prompt, query, or conversation title,
  • Name of person who queried

41 of 53

Cite Your Sources!... and Tools

APA citation:

Hepler, R. and OpenAI, (2023). "[Chat title]", conversation with [tool name] [Large Language/Image Model]  ([version information]). Generated on [date]. [shareable link to the chat, if possible].

For example, I would put 

Hepler, R., and OpenAI. (2023). "Balrogs might have wings", online conversation with ChatGPT [Large Language Model] (August 3 Version). Generated on August 22, 2023. https://chat.openai.com/share/15d75e9f-16d3-4ebf-81b8-f675528ed267.

The MLA analogue would be:

Hepler, R. and OpenAI. "Balrogs Might Have Wings." Conversation with ChatGPT. August 3 Version, 22 Aug. 2023, chat.openai.com/share/15d75e9f-16d3-4ebf-81b8-f675528ed267.

42 of 53

Questions?

“In some cases, we learn more by looking for the answer to a question and not finding it than we do from learning the answer itself.” - Dallben, The Book of Three

43 of 53

Three Cs of Generative AI

Circumspection

what hazards (moral, ethical, educational) should I manage?

44 of 53

Important Aspects of

OpenAI’s Terms of Use

  • “Input and Output are collectively ‘Content.’ You are responsible for Content, including ensuring that it does not violate any applicable law or these Terms. 
  • You (a) retain your ownership rights in Input and (b) own the Output. We hereby assign to you all our right, title, and interest, if any, in and to Output.” 
  •  https://openai.com/policies/terms-of-use
  • To have a more privacy-conscious session with ChatGPT, you can use the platform.openai.com Playground iteration of ChatGPT. 
  • Or, you can “opt out” and still have access to your chat history at https://privacy.openai.com/policies

45 of 53

Generalizability of Data Protection Practices and Safeguards

Be extra vigilant; put existing principles to use in new practices.

We deliberately and directly give our private and confidential data to generative AI tools

Even with delete buttons our data is still sold and retained.

In some ways, Generative AI is no more dangerous than institutions that have our data through other means.

46 of 53

Considerations for Using AI in the Workplace

  • Transparent Data Collection
  • Informed Consent
  • Data Minimization

"Business Man" by Direct Media is marked with CC0 1.0.

47 of 53

Specific Recommendations for Safeguarding Privacy

  • Avoid explicit mention of private details. 
  • Avoid mentioning any generic factors of your life.
  • Implement and enforce strict access and use controls.
  • Have family members or friends read over your prompts.
  • Never assume that opting out will protect privacy.
  • Never put individual peoples’ attributes into their prompts.

48 of 53

Specific Recommendations for Safeguarding Confidentiality

Ensure that staff and contractors are sure as to what data they can input into their prompts.

Enforce a “least privileged access” model. 

Use human content moderation. 

Anonymize all data, including advertisements, business plans, marketing plans, etc.

49 of 53

Discussion

50 of 53

Argento, Z. (2023, August 9). Data protection issues for employers to consider when using generative AI. https://iapp.org/news/a/data-protection-issues-for-employers-to-consider-when-using-generative-ai/ 

Citron, D. K., & Solove, D. J. (2021, February 18). Privacy harms. SSRN. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3782222.

Falconer, S. (2023, October 23). Privacy in the age of Generative AI. Stack Overflow. https://stackoverflow.blog/2023/10/23/privacy-in-the-age-of-generative-ai/  

Hepler, R. and OpenAI (2023). “AI Ethics in Education,” online conversation with GPT 4 [Large Language Model]. Generated on December 11, 2023. https://chat.openai.com/share/884457f6-96f4-404b-8a2b-c8bb6c8ad041

Mancuso, D. (2023, July 17). Privacy & Cybersecurity. Privacy Cybersecurity. https://cybersecurity.illinois.edu/privacy-considerations-for-generative-ai/.

OpenAI. (2023, November 14). Terms of use. https://openai.com/policies/terms-of-use 

Rose, R. (2023, April 10). Ethical considerations. ChatGPT in Higher Education. https://unf.pressbooks.pub/chatgptinhighereducation/chapter/chapter-2/ 

Yousefzadeh, R., & Cao, X. (2022, January 27). To what extent should we trust AI models when they extrapolate?. arXiv.org. https://arxiv.org/abs/2201.11260 

References

51 of 53

Acknowledgments

Thanks to Nathan Hunter for his excellent book, The Art of Prompt Engineering with ChatGPT, which helped me become proficient in many styles of prompts and understand the contexts in which they are most appropriate. 

The Art of Prompt Engineering with ChatGPT, by Nathan Hunter. 

Co-Intelligence, by Ethan Mollick

Hepler Consulting Website

Hepler Consulting LinkedIn

52 of 53

Hepler Consulting

LinkedIn

heplerconsulting.com

CollaborAItion

Blog

BlueSky

53 of 53

THANK YOU