1 of 70

Generative KI, LLMs und GPT bei digitalen Editionen

DHd2024. Universität Passau�27.02.2024

Christopher Pollin | Alexander Czmiel |

Torsten Roeder | Torsten Schaßan |

Patrick Sahle | Martina Scholger |

Franz Fischer | Stefan Dumont |

Georg Vogeler | Christiane Fritze�Institut für Dokumentologie und Editorik

2 of 70

  • Gerrit Brüning, Felix Schenke: Umwandlung von tabellarischen Daten in TEI-XML mithilfe von Oxygen AI Positron �
  • Bastian Politycki: Anwendung generativer KI zur Digitalisierung gedruckter Editionen am Beispiel der Sammlung Schweizerischer Rechtsquellen�
  • Carina Geldhauser, Ipek Tuncel: Halbautomatische Annotierung antiker Handschriften�
  • Kay-Michael Würzner, Robert Sachunsky: Korrektur & (De-)Normalisierung historischer Volltexte�
  • Yannic Bracke: LLM-basierte Normalisierung historischer Schreibweisen mit transnormer�
  • Pia Schwarz, Florian Barth, Lennart Keller: Klassifikation und Linking von Entitäten. Spezifischer Klassifikator vs. Large Language Model�
  • Tarjia Alam Nisha, Franziska Pannach, Jörg Wettlaufer: Itinerare erkennen in Reiseberichten. Auszeichnung von Orts- und Personennamen zur Etablierung von Itineraren in Reiseberichten des 19. Jahrhunderts.�
  • Dominic Fischer, Martin Volk, Patricia Scheurer, Phillip Ströbel: LLMs for Bullinger Digital�
  • Jacob Möhrke, Sandra Balck, Anna Ananieva: Zum Einsatz von GPT-4 für NER: Ein Experiment anhand eines historischen Reisetextes�
  • Nina Claudia Rastinger: Informationsextraktion aus frühneuzeitlichen Ankunftslisten – das Projekt „Visiting Vienna“ als Fallstudie zur Named Entity Recognition mit GPT-3.5

3 of 70

Agenda

Vormittags-Session (09:00 – 12:45)

  • 09:00 – 10:00: Einführung (Pollin)
  • 10:00 – 10:45: Experiment 1 (Brüning-Schenke)
  • 10:45 – 11:00: Pause
  • 11:00 – 11:45: Experiment 2 (Politycki)
  • 11:45 – 12:45: Experiment 3-5 (Geldhauser-Tuncel & Würzner-Sachunsky & Bracke)

Mittagspause (12:45 – 14:15)

Nachmittags-Session (14:15 – 17:30)

  • 14:15 – 15:00: Experiment 6 (Schwarz-Barth-Keller)
  • 15:00 – 15:45: Experiment 7 (Nisha-Pannach-Wettlaufer)
  • 15:45 – 16:00: Pause
  • 16:00 – 17:00: Experiment 8-9 �(Fischer-Volk-Scheuer-Ströbl & Möhrke-Balck-Ananieva & Rastinger)
  • 17:00 – 17:30: Abschlussdiskussion (eingeleitet von Sahle et al)

Themenschwerpunkte �(ungefähr nach Editions-Workflow):

  1. Überlieferungsdokumentation
  2. Retro-nachbearbeiten
  3. Textherstellung (Transkription, OCR Cleanup, Markup-Erzeugung)

  • Normalisierung
  • NER
  • Annotation
  • Übersetzung & Text-Zusammenfassung

4 of 70

Von “bad prompts” mit ChatGPT-3.5 zu Workflows mit GPT-4 Agenten, unterstützt durch Custom GPTs und GPT-Vision: Eine Analyse am Beispiel eines Briefes

Friedrich August Otto Benndorf an Hugo Schuchardt (02-00932). Wien, 14. 02. 1879. Hrsg. von Hubert Szemethy (2022). In: Bernhard Hurch (Hrsg.): Hugo Schuchardt Archiv. Online unter https://gams.uni-graz.at/o:hsa.letter.7711, abgerufen am 07. 06. 2023. Handle: hdl.handle.net/11471/518.10.1.7711.

5 of 70

GPT-3.5

TEI XML Brief erstellen. January 29, 2024. GPT-3.5. https://chat.openai.com/share/e/068b765c-2464-49e1-a6ec-2c30fb7e8808

Friedrich August Otto Benndorf an Hugo Schuchardt (02-00932). Wien, 14. 02. 1879. Hrsg. von Hubert Szemethy (2022). In: Bernhard Hurch (Hrsg.): Hugo Schuchardt Archiv. Online unter https://gams.uni-graz.at/o:hsa.letter.7711, abgerufen am 07. 06. 2023. Handle: hdl.handle.net/11471/518.10.1.7711.

  • GPT-3.5 kann wohlgeformtes XML erzeugen�
  • Selten oder nie valides TEI XML

  • … eigentlich ist das Ergebnis in diesem Fall gar nicht so schlecht�
  • “Bad Prompting”

6 of 70

Prompt Engineering matters!

Bsharat, Sondos Mahmoud, Aidar Myrzakhan, and Zhiqiang Shen. “Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4.” arXiv, December 26, 2023. https://doi.org/10.48550/arXiv.2312.16171.

Nori, Harsha, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, et al. “Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine.” arXiv, November 27, 2023. https://doi.org/10.48550/arXiv.2311.16452.

Verbessern bei GPT-4 (laut Studie) …

  • 75-85% die Korrektheit von Antworten
  • 40-80% die Qualität (im weiteren Sinne) von Antworten
  • GPT-4 mit Prompting übertrifft (oft) fine tuned LLMs

7 of 70

Prompt Engineering ist auch absurd: “Llama2-70B ist ein Trekkie”

Battle, Rick, and Teja Gollapudi. “The Unreasonable Effectiveness of Eccentric Automatic Prompts.” arXiv, February 20, 2024. https://doi.org/10.48550/arXiv.2402.10949.

8 of 70

Prompt Engineering

You will act as a skilled expert automaton that is proficient in transforming unstructured text, specifically multilingual letters from or to Hugo Schuchardt (1842-1927), into well-formed TEI XML. Analyze the provided text based on the mapping rules I have shared and then execute the transformation to produce TEI XML, ensuring you adhere to the guidelines and only annotate if certain.

Mapping rules:

* <div> Entire letter

* <pb> Marks page breaks e.g. "|{n}|", multiple appearance possible, always as child of <div>

* <dateline> Date/time reference of the letter

* <date> in <dateline>

* <opener> Opening of the letter

* <closer> Closing of the letter

* <salute> Salutations within the letter

* <lb> Line breaks

* <signed> Signature section

* <postscript> Represents a postscript

* <bibl> Contains bibliographical references

* <p> Paragraphs

* <persName> Person

* <placeName> Place

* <orgName> Organisation

* <date> Dates; when={YYYY-MM-DD}

* <term> Languages

* <foreign> Words in the context of discussing the linguistic phenomenon

Guidelines:

* Strictly follow mapping rules

* Preserve the original text

* Produce well-formed TEI XML according to TEI standards

* Return the <div> only

* Annotate only when appropriate

* Preserve complexity of output

* Compact XML without any whitespace or indentation��Brief von Friedrich August Otto Benndorf an Hugo Schuchardt:�´´´�{text}�´´´

This is very important for my career!

Persona Modelling

Context�

Tasks�

Spezifität + ~”Few-Shot Prompting”�

Emotional prompting

Bsharat, Sondos Mahmoud, Aidar Myrzakhan, and Zhiqiang Shen. “Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4.” arXiv, December 26, 2023. https://doi.org/10.48550/arXiv.2312.16171.

Li, Cheng, Jindong Wang, Yixuan Zhang, Kaijie Zhu, Wenxin Hou, Jianxun Lian, Fang Luo, Qiang Yang, and Xing Xie. “Large Language Models Understand and Can Be Enhanced by Emotional Stimuli.” arXiv, November 12, 2023. https://doi.org/10.48550/arXiv.2307.11760.

Pollin, C. (2024). Workshopreihe "Angewandte Generative KI in den (digitalen) Geisteswissenschaften" (v1.1.0). Zenodo. https://chpollin.github.io/GM-DH/

9 of 70

GPT-3.5 + �Prompt Engineering

  • Das Ergebnis ist reicher an Annotationen und Normalisierungen�
  • Nicht valides TEI�
  • GPT-3.5 weist wenig “Reasoning” Kapazitäten auf

10 of 70

GPT-4 + �Prompt Engineering

  • GPT-4 weist hingegen viel bessere Reasoning Kapazitäten auf�
  • Es hält sich viel eher an Regeln:
    • Preserve complexity of output
    • Return the <div> only
    • Produce well-formed TEI XML according to TEI standards

  • Nur <opener> macht die Validität kaputt�
  • <foreign> passt nicht: kanns nicht erklären!
  • Aber: “8. D. M.” und “11.” wurden (fast) richtig gefunden und normalisiert.

11 of 70

Workflow: GPT-4 + Prompt Engineering + API

Unstructured Text

GPT API

TEI XML

Validation

fail

System Prompt

n Briefe werden mittels eines Python Scripts nach TEI XML transformiert

(optional) preprocessing

LLM

no LLM

Output TEI XML

postprocessing

12 of 70

Workflow: GPT-4 + Prompt Engineering + API + “Editor in the Loop”

12

Unstructured Text

GPT API

TEI XML

Validation

fail

System Prompt

Editor (Human)

n Briefe werden mittels eines Python Scripts nach TEI XML transformiert

(optional) preprocessing

LLM

no LLM

human

postprocessing

Final TEI XML

postprocessing

13 of 70

GPT-4 + Prompting + Knowledge + RAG (Custom GPT)

https://chat.openai.com/g/g-FEUt7Fq48-teicrafter

14 of 70

Custom GPT: teiModeler

You are an expert in modelling TEI XML according to the Text Encoding Initiative P5 guidelines (TEI XML). Your main objective is to find the best text model for a given text using TEI XML.

You will do the following:

* Analyse the text very carefully and define the type of text.

* Discuss all text phenomena in detail.

* Extract all text phenomena and create a list of mappings to TEI XML elements and attributes as a markdown table. All existing elements are listed in TEI Elements.md and all existing attributes are listed in TEI Attributes.md. You must use these elements and attributes.

* Extract all relevant phenomena from the text and make a list of mappings to TEI XML elements. Discuss the mapping in detail.

* Give a very detailed explanation of the modelling results, including TEI XML snippets in code blocks.

* Give 2 different ways of modelling.

* Ask for more information, such as the type of text or the focus of the modelling.

Rules:

* Ignore parent elements such as <TEI>, <body>, <text>, <teiHeader>.

* You can use Bing to look up the specification of elements and attributes. This is the URL for the <seg> element: https://www.tei-c.org/release/doc/tei-p5-doc/en/html/ref-seg.html

* NEVER change the input text

* ALWAYS create valid and well-formed TEI XML.

Always end with:

´´´

This is just one approach to modelling. Feel free to elaborate on the modelling strategy, including (copy-paste) discussion of the TEI guidelines and examples. Keep in mind that my answers may contain inaccuracies or fabricated information. Feel free to ask me any questions!

´´´

Let's work on this step by step! This is very important for my career!

Instruction

Knowledge

* TEI Attributes.md�* TEI Elements.md�* Attribute Classes.md

15 of 70

teiModeler Beispiel 1/2

16 of 70

teiModeler Beispiel 2/2: valides TEI

17 of 70

Workflow: GPT-4 + Prompt Engineering + API + “Editor in the Loop” + �Assistance API

Actions

RAG

Unstructured Text

teiCrafter

GPT API

teiModeler

Editor (Human)

Feedback

Deterministic Validation�(e.g. Schema)

teiVerifier

Actions

RAG

Assistance GPT API

Assistance GPT API

Final TEI XML

UI for verifying and annotating

Iterations

18 of 70

Workflow: GPT-4 + Prompt Engineering + API + Assistance API + “Editor in the Loop” + �Multimodalität GPT-Vision

Actions

RAG

Unstructured Text

teiCrafter

GPT API

TEI Modeler

Editor (Human)

Feedback

Deterministic Validation�(e.g. Schema)

teiVerifier

Actions

RAG

Assistance GPT API

Assistance GPT API

Final TEI XML

UI for verifying and annotating

GPT-4 Vision

GPT API

Contextual information about the digital facsimile

19 of 70

Multi Agent TEI XML Creation Piplin

Alle moderenrn technicken gemeinsam abbilden, aber sagendass ich ihn nicht gebaut habe, aber das müsste klappen und ist als AI Engineering komplex! Darum baut es auch keiner. Oder es ist noch nicht da.

Text → tei modeller mit veryfier and reasoning (o1) o1 übergibt aufgaben auf andere task inklusive einem umfangreichen reasoning. Also o1 prompted die anderen spezialisierten modelle die agents führen in gruppen ihre tasks druch und aggregieren ihre gesammelten ergebnisse, es ist alles mega touer in token. Wir m+üssen erklären was toekn sind und was api kostes.

20 of 70

Agents

AutoGen: “Build LLM applications via multiple agents”

Wang, Guanzhi, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, und Anima Anandkumar. „Voyager: An Open-Ended Embodied Agent with Large Language Models“, 25. Mai 2023. https://arxiv.org/abs/2305.16291v2.

Wu, Qingyun, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, et al. “AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation,” August 16, 2023. https://arxiv.org/abs/2308.08155v2.

AutoGen. https://www.microsoft.com/en-us/research/project/autogen/

21 of 70

Workflow: GPT-4 + Prompt Engineering + API + Assistance API + “Editor in the Loop” + Multimodalität GPT-Vision +

Agents

Skill: fetch_tei_xml

Skill: reconcile-entities

22 of 70

AI Buzzwords: �weil man es alleine nicht schafft sich alles anzuschauen!

  • LangChain
  • Function Calling
  • AutoGPT
  • (Advanced) RAG
  • Vektordatenbanken
  • Fine-Tuning
  • Let's verify step by step

23 of 70

Ausblick

  • “GPT-2 couldn't do very much, GPT-3 could do more, GPT 4 could do a lot more, GPT 5 [in Training] will be able to do a lot lot more” (Sam Altman, OpenAI)
  • Google Gemini 1.5 hat ein Context Window von 1.000.000 (multimodalen!) Token (30K Zeilen Code, 700 Seiten Text, 11 Stunden Audio, … )
  • AlphaCode 2 (Google) erreicht extrem gute Ergebnisse beim Programmieren
  • OpenAI hat zwei Patente mit dem Ziel der “Automatisierung des Programmierens”
  • Autonome Agenten zeichnen sich ab
  • Exponentielle Entwicklung im AI Bereich!?�Was macht das mit den DH und DigEd?

Sam Altman Just Revealed NEW DETAILS About GPT-5 In Spicy 🌶️ Interview. https://www.youtube.com/watch?v=RYg5Mz4_tf8 �“LLMs Will Make Programming Useless In 10 Years”. https://youtu.be/ZV6Sz42l0hY?si=cMOZ02r6tLBqZtTD �AlphaCode 2 Technical Report. 06.12.2023. https://storage.googleapis.com/deepmind-media/AlphaCode2/AlphaCode2_Tech_Report.pdf �AI Explained. Gemini Full Breakdown + AlphaCode 2 Bombshell. https://www.youtube.com/watch?v=toShbNUGAyo&t=1s �GPT-5: Everything You Need to Know So Far. https://www.youtube.com/watch?v=Zc03IYnnuIA �OpenAI Patente: https://www.freepatentsonline.com/y2024/0020096.html, https://www.freepatentsonline.com/y2024/0020116.html �Gemini 1.5 and The Biggest Night in AI. https://www.youtube.com/watch?v=Cs6pe8o7XY8&list=PLaHADNRco7n3GKVUD8mAc36pXQ5pnJQVL&index=8

24 of 70

Ressourcen

Pollin, C. (2024). Workshopreihe "Angewandte Generative KI in den (digitalen) Geisteswissenschaften" (v1.1.0). Zenodo. https://doi.org/10.5281/zenodo.10647754

Applied Generative AI in Digital Humanities. YouTube Playlist. https://youtube.com/playlist?list=PLaHADNRco7n3GKVUD8mAc36pXQ5pnJQVL&si=sHC_ColVJ0J9vpHx

AGKI-DH. Zotero Group. https://www.zotero.org/groups/5319178/agki-dh

25 of 70

Umwandlung von tabellarischen Daten in TEI-XML mithilfe von �Oxygen AI Positron

Gerrit Brüning, Felix Schenke

https://docs.google.com/presentation/d/1o6vwfL1IZSSUgYnRIP4J2qm3t9mLJmH4/edit?usp=sharing&ouid=115137275675286426842&rtpof=true&sd=true

Experiment 1

26 of 70

Anwendung generativer KI zur

Digitalisierung gedruckter

Editionen am Beispiel der

Sammlung Schweizerischer

Rechtsquellen

Bastian Politycki

https://drive.google.com/file/d/1kJs0NUkrj28qM5UeaFFWXDhIZGhIHdaU/view?usp=drive_link

Experiment 2

27 of 70

Halbautomatische Annotierung

antiker Handschriften

Carina Geldhauser

Ipek Tuncel

https://drive.google.com/file/d/1Dn0PxngZ_XuHPNiAu_7TZYki9f15c-XE/view?usp=drive_link

Experiment 3

28 of 70

Korrektur & (De-)Normalisierung historischer Volltexte

Kay-Michael Würzner

Robert Sachunsky

https://docs.google.com/presentation/d/1rOWYKXlxr8QPZA43dn5lvJXgrtWOgL4tLRkwf3VI3NA/edit?usp=sharing

Experiment 4

29 of 70

LLM-basierte Normalisierung historischer Schreibweisen mit transnormer

Yannic Bracke

https://drive.google.com/file/d/19cgdVF5vQdaPhZjzf57Pw7cIMKT-e6R9/view?usp=sharing

Experiment 5

30 of 70

Klassifikation und Linking von Entitäten. Spezifischer Klassifikator vs. Large Language Model

Pia Schwarz

Florian Barth

Lennart Keller

https://pad.gwdg.de/p/1kJ6AiaJO#/

Experiment 6

31 of 70

Itinerare erkennen in Reiseberichten. Auszeichnung von

Orts- und Personennamen zur Etablierung von Itineraren in Reiseberichten des 19. Jahrhunderts.

Tarjia Alam Nisha

Franziska Pannach

Jörg Wettlaufer

https://drive.google.com/file/d/1N-p2JdJOMz2CWyMtEl4emh84LoHIJaVz/view?usp=sharing

Experiment 7

32 of 70

Experiment 8

33 of 70

Zum Einsatz von GPT-4 für NER:

Ein Experiment anhand eines

historischen Reisetextes

Jacob Möhrke

Sandra Balck

Anna Ananieva

https://drive.google.com/file/d/1ODfrr9mcPI3sEfr6YsQiLIuWfFCCXk3s/view?usp=sharing

Experiment 9

34 of 70

Informationsextraktion aus frühneuzeitlichen Ankunftslisten – das Projekt „Visiting Vienna“ als Fallstudie zur Named Entity Recognition mit GPT-3.5

Nina Claudia Rastinger

Experiment 10

35 of 70

Abschlussdiskussion

  1. Wrap up
  2. Diskussion

36 of 70

Abschlussdiskussion

wrap up / Einordnungsversuch

  • Produktiv 1: Der Dschungel der Möglichkeiten
  • Produktiv 2: Workflow-Orchestrationen
  • Reflexiv 1: Stärken und Schwächen der KI
  • Reflexiv 2: Fokusverschiebungen?

37 of 70

wrap up / Einordnungsversuch

Produktiv 1: Der Dschungel der Möglichkeiten

  • Die verschiedenen LLMs
  • Generische Anwendungen
  • Spezialisierte Tools, Add-Ons
  • Prompt Engineering
  • RAG et al.
  • Eigene Entwicklungen: Vektor-DB, Feintuning, Training, CustomGPT, actions et al.
  • Was steht vor der Tür?

38 of 70

wrap up / Einordnungsversuch

Produktiv 2: Workflow-Orchestrationen

  • An welcher Stelle KI einsetzen / für welche Aufgabe? Für welche nicht?
  • Generische Anwendungen
  • Spezialisierte Tools, Add-Ons
  • Eigene Entwicklungen, Fine-Tunings
  • Zusammenspiel mit anderen Komponenten (pre-, post-, etc.)
  • Evaluation und Qualitätssicherung
    • Schlechte Lösung? Replizierbarkeit, Benchmarking
  • Effizienzen, Kosten-Nutzen-Situation aktuell
  • Ausblick?

39 of 70

wrap up / Einordnungsversuch

Reflexiv 1: Stärken und Schwächen

  • Übersetzungen, Zusammenfassungen, Textverbesserungen
  • Semantische Zusammenhänge, common sense
  • Verstehen, Heterogenität, Unschärfe, Komplexität
  • Multimodalität

  • Qualität → Relation zu Erwartungen, Nutzbarkeit in Workflows

  • Kontextlänge
  • Aktionen
  • Vollständigkeit, Präzision
  • Zuverlässigkeit, Halluzination
  • Fakten und Regeln

Stärken

Schwächen

40 of 70

wrap up / Einordnungsversuch

Reflexiv 2: Fokusverschiebungen

  • Neues Tool für alte Aufgaben?
    • Was ist unser “Standardtool”?
  • Neue Zielstellungen?
    • Veränderte Epistemologie

41 of 70

Abschlussdiskussion

  • Wo stehen wir?
  • Wie geht es weiter?

42 of 70

Anhang

43 of 70

Handwriting Text Recognition (HTR): Bereinigen des GPT-4 Vision + Transkribus Ergebnisses (“No Human in the Loop”)

44 of 70

Hype?! �Es geht erst richtig los!

AI Explained. 4 Reasons AI in 2024 is On An Exponential: Data, Mamba, and More. https://www.youtube.com/watch?v=Xq-QEd1jpKk&t=298s

Phi-2: The surprising power of small language models. https://www.microsoft.com/en-us/research/blog/phi-2-the-surprising-power-of-small-language-models/

Mamba: Linear-Time Sequence Modeling with Selective State Spaces. https://arxiv.org/abs/2312.00752.

https://www.etched.com

Applied Generative AI in Digital Humanities. https://youtube.com/playlist?list=PLaHADNRco7n3GKVUD8mAc36pXQ5pnJQVL&si=_z7jlFOkbK9ur3bu

GPT-5-Tier LLM

“Echte” KI-UI und Tools

Autonome Agenten

Multimodalität, Embodiment, synthetische Daten, Mamba, etched, ...

45 of 70

Large Language Models (LLM)

Vs. �Digitale Edition

Edition: viel Information zu einem exakten? Text

“Gestalt” von Text

“LLM are like having a Zip-File of the internet”

* Midjourney: https://s.mj.run/g7Mm_h0ZH9w hyper realistic and sureal gigantic yellow folder with a zipper, like a desktop icon, ultra detailed, salvador dali desert background, landsacape --ar 16:9 --v 6.0 --style raw --stylize 800 �* magnific.ai

Andrej Karpathy. [1hr Talk] Intro to Large Language Models. https://www.youtube.com/watch?v=zjkBMFhNj_g&list=WL&index=16

46 of 70

Transformer-Architektur

46

Andrej Karpathy. [1hr Talk] Intro to Large Language Models. https://www.youtube.com/watch?v=zjkBMFhNj_g&list=WL&index=16

47 of 70

47

Die Bibliothek von

Babel

Infinite Monkey�Theorem

Stochastic Parrot

DALL-E 3: A triptych where each section is visually distinct. Section 1: An ancient library filled with tall wooden bookshelves, dusty tomes, and dim candlelight, invoking a sense of age and wisdom. Section 2: Multiple monkeys at individual typewriters in a surreal, abstract space, with papers flying around, suggesting chaotic creativity. Section 3: A single parrot speaking into a microphone, with a background of digital screens showing strings of text and code, representing the voice output of text generated by algorithms.�magnific.ai:

Bender, Emily M., Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜.” In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–23. FAccT ’21. New York, NY, USA: Association for Computing Machinery, 2021. https://doi.org/10.1145/3442188.3445922.

48 of 70

Token & Embedding

48

Token

  • Teile von Text und Input für LLM
  • 1 Token entspricht ~4 Zeichen englischen Standardtextes 100 Token ~= 75 Wörter.

Embedding

  • Darstellung des Textes als Zahlen in einem mehrdimensionalen Vektorraum.
  • Stellt die "Bedeutung" des Textes im LLM dar.

A minimalist and artistic infographic showing geometric, stylized figures of a dog and cat adjacent to each other on a subtly illuminated 3-dimensional vector space grid with the labels 'dog' and 'cat' in a clear, professional font. At a significant distance, a stone with a sad face emoticon is placed, isolated from the animals, with the label 'stone'. The color palette is muted and sophisticated, enhancing the professional aesthetic.

49 of 70

Prompt & Prompt Engineering

49

Prompt �ist die natürlichsprachliche Eingabe, die dem Modell (z. B. LLM) zur Verfügung gestellt wird und auf die das Modell reagiert.

Prompt Engineering �ist der Prozess des Entwerfens, Verfeinerns und Optimierens von Prompts, um die Absicht der User*innen effektiv an ein LLM zu kommunizieren.

* Midjourney: https://s.mj.run/tcdb6wtkzj4 engineer wizard, in front of computer, workshop, comic style, Working with tools, welding --ar 32:18

* magnific.ai

50 of 70

Prompt & Prompt Engineering

50

  • Persona Modelling: � “You are an expert…
  • Context Information
  • Chain of Thought: � “Let's think step by step
  • Output vorgeben: � “markdown table

* Midjourney: https://s.mj.run/tcdb6wtkzj4 engineer wizard, in front of computer, workshop, comic style, Working with tools, welding --ar 32:18

* magnific.ai

51 of 70

Prompt Engineering Prinzipien

  • Spezifität und Klarheit�Die Aufforderungen sollten klar und eindeutig formuliert sein, um ungenaue oder unerwünschte Ergebnisse zu vermeiden.�
  • Zeit zum “Nachdenken” einplanen�Es ist wichtig LLMs genügend Zeit zu geben, um Informationen zu verarbeiten.�
  • Kontext und Beispiele verwenden�Die Bereitstellung von Kontext und Beispielen kann die Qualität und Relevanz der Antworten des Modells verbessern.�
  • Iterativer Ansatz�Die Entwicklung von Prompts erfordert oft wiederholte Anpassungen, daher ist es wichtig, eine offene Haltung und die Bereitschaft zu bewahren, die Prompts auf der Grundlage der erhaltenen Antworten zu verfeinern.

  • Verstehen der Fähigkeiten von GPTDas Modell eignet sich hervorragend zum Zusammenfassen, zum Ableiten von Informationen, zum Konvertieren von Daten in verschiedene Formate, zum Generieren von Ideen, etc. …

51

  • Explizite Einschränkungen verwendenDas Einfügen klarer Grenzen oder Richtlinien in der Prompt kann helfen zu kontrollieren, wie das Modell reagiert.
  • Vermeiden Sie Überlastung

Zu komplexe oder zu viele Aufgaben auf einmal können für das Modell problematisch sein und zu ungenauen oder unvollständigen Antworten führen. Oft ist es ratsam, solche Anforderungen in überschaubare Segmente aufzuteilen.�

  • Multimodale BetrachtungKI Modelle nicht mehr nur textbasiert. Ein weiteres Prinzip kann sein, zu überlegen, wie Prompts in multimodalen Modellen (Kombination von Text, Bild, Audio etc.) funktionieren.

52 of 70

Custom Instructions

Eine Custom Instruction ist eine System Prompt.

Anweisungen, die das Modell berücksichtigt, bevor es eine Antwort generiert.

Sie beeinflussen:

  • Wie ist die Antwort: Detailgrad, Ton, Stil, …
  • Wer generiert Text für wen: Persona Modeling, Zielgruppe, …
  • Weitere Regeln: Verwende X, …

52

You are an expert in world history, knowledgeable about different eras, civilizations, and significant events. Provide detailed historical context and explanations when answering questions. Be as informative as possible, while keeping your responses engaging and accessible.

53 of 70

Context Window

Im Zusammenhang LLMs bezieht sich ein Context Window auf die Textmenge (in Form von Tokens), die das Modell bei der Erzeugung von Antworten gleichzeitig berücksichtigen kann.

Dieses Fenster bestimmt die Menge an Informationen, die das Modell zum Prozessieren (Simulation von Reasoning) und Generieren jedes Teils seiner Ausgabe verwenden kann.

53

54 of 70

Context Window: “Lost in the Middle”

LLM funktionieren am besten, wenn die wichtigen Informationen am Anfang oder Ende des Eingabekontextes stehen.

Es gibt einen signifikanten Leistungsabfall, wenn Modelle Informationen verarbeiten müssen, die in der Mitte von langen Kontexten platziert sind.

Dieses Problem besteht auch bei Modellen, die speziell für die Verarbeitung längerer Kontexte entwickelt wurden.

Grob gesagt:

  • Wichtigstes am Anfang und am Ende!
  • Weniger Tokens ist besser!

54

Liu, Nelson F., Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. “Lost in the Middle: How Language Models Use Long Contexts.” arXiv, November 20, 2023. http://arxiv.org/abs/2307.03172.

55 of 70

Context Window: “Lost in the Middle”

55

https://twitter.com/GregKamradt/status/1722386725635580292

“Needle-in-a-haystack experiments”.

Ivgi, Maor, Uri Shaham, and Jonathan Berant. “Efficient Long-Text Understanding with Short-Text Models.” Transactions of the Association for Computational Linguistics 11 (2023): 284–99. https://doi.org/10.1162/tacl_a_00547.

56 of 70

Zero-Shot Prompting

Few-Shot Prompting

Classify the text into neutral, negative or positive.

Text: I think the vacation is okay.

Sentiment:

A "whatpu" is a small, furry animal native to Tanzania. An example of a sentence that uses the word whatpu is:

We were traveling in Africa and we saw these very cute whatpus.

To do a "farduddle" means to jump up and down really fast. An example of a sentence that uses the word farduddle is:

56

57 of 70

Chain-of-Thought Prompting

57

Wei, Jason, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, und Denny Zhou. „Chain-of-Thought Prompting Elicits Reasoning in Large Language Models“. arXiv, 10. Januar 2023. https://doi.org/10.48550/arXiv.2201.11903.

Yao, Shunyu, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. ‘Tree of Thoughts: Deliberate Problem Solving with Large Language Models’. arXiv, 17 May 2023. https://doi.org/10.48550/arXiv.2305.10601.

58 of 70

Lets verify step by step

Nicht eine antwort egenrieren, sondern 1000 von antworten generierne und ein zweites modell überprüft was richtig ist

https://www.youtube.com/watch?v=Zc03IYnnuIA

Alpha Code Technical Report

Lightman, Hunter, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. “Let’s Verify Step by Step.” arXiv, May 31, 2023. https://doi.org/10.48550/arXiv.2305.20050.

59 of 70

OCR Cleaning / Text reparieren / Vervollständigen

Cao, Qi, Takeshi Kojima, Yutaka Matsuo, and Yusuke Iwasawa. “Unnatural Error Correction: GPT-4 Can Almost Perfectly Handle Unnatural Scrambled Text.” arXiv, November 30, 2023. https://doi.org/10.48550/arXiv.2311.18805.

60 of 70

Custom Instructions

Eine Custom Instruction ist eine System Prompt.

Anweisungen, die das Modell berücksichtigt, bevor es eine Antwort generiert.

Sie beeinflussen:

  • Wie ist die Antwort: Detailgrad, Ton, Stil, …
  • Wer generiert Text für wen: Persona Modeling, Zielgruppe, …
  • Weitere Regeln: Verwende X, …

60

You are an expert in world history, knowledgeable about different eras, civilizations, and significant events. Provide detailed historical context and explanations when answering questions. Be as informative as possible, while keeping your responses engaging and accessible.

61 of 70

Keine Custom Instruction

“Expert in World History”- Custom Instruction

Custom Instructions

61

62 of 70

Custom Instructions

62

63 of 70

“Universale” Custom Instruction zur grundlegenden Verbesserung von GPT-4

63

This is relevant to EVERY prompt I ask.

Never tell me “As a large language model…” or “As an artificial intelligence…”

I already know you are an LLM. Just tell me the answer.

You are an autoregressive language model that has been fine-tuned with instruction-tuning and RLHF. You carefully provide accurate, factual, thoughtful, nuanced answers, and are brilliant at reasoning. If you think there might not be a correct answer, you say so.

Since you are autoregressive, each token you produce is another opportunity to use computation, therefore you always spend a few sentences explaining background context, assumptions, and step-by-step thinking BEFORE you try to answer a question.

Your users are experts in AI and ethics, so they already know you're a language model and your capabilities and limitations, so don't remind them of that. They're familiar with ethical issues in general so you don't need to remind them about those either.

Don't be verbose in your answers, but do provide details and examples where it might help the explanation.

64 of 70

Custom GPTs

64

Midjourney: https://s.mj.run/7WRa7TAMzck award winning illustration, researcher in an academic setting, full body, multiple semi-transparent holographic displays, university setting, muted academic tones, balance of realism and illustration, Gustav Klimt style, --ar 16:9 --style ZEfVSLa1�Zoom Out:

Variations (Region): research data, network, graph, nodes, historical text, digital humanities, wirting, documents --ar 16:9

https://magnific.ai/

Custom Instructions

Tools: �Browsing, DALL·E, Code Interpreter

Custom Actions

Knowledge Base

65 of 70

Warum Custom GPTs

  • Halluzinationen Reduzieren
  • Personalisierung von GPT
    • Eigene Daten und eigenes Wissen verwendet
    • User Interaction anpassen
    • Eigene Scripte und APIs dranhängen
  • Optimierung von Workflows
  • Niederschwellige Entwicklung

65

66 of 70

Custom GPTs: GPT Store & Consensus.ai

66

67 of 70

Custom GPTs erzeugen

67

GPT Builder

Prompting

68 of 70

Custom GPTs: Consensus.ai

68

69 of 70

Consensus.ai

69

Juggling Roles, Experiencing Dilemmas: The Challenges of SSH Scholars in Public Engagement (2021) by J. Schuijer et al.

This paper explores the new roles and challenges faced by Social Science and Humanities (SSH) scholars in public engagement, especially in the context of emerging technologies like nanotechnology. Read more.

It is Essential to Connect: Evaluating a Science Communication Boot Camp (2022) by Krista Longtin et al.

This study evaluates the effectiveness of a Science Communication Boot Camp in improving participants' communication skills and willingness to engage with the public. Read more.

Integrative Approaches to Dispersing Science: A Case Study of March Mammal Madness (2021) by C. E. G. Amorim et al. This paper discusses the importance of public engagement as a pillar of scientific scholarship and the challenges faced in science communication. Read more.

I am interested in public engagement.

Please list the top publications on this topic with a focus on science communication in the humanities. All publications must be younger than 2020 and in english or german.

70 of 70

Consensus.ai

70