1 of 21

Generating Laughter�Evaluating the Success of LLMs for Comedic Purposes

Erin Mikail Staples (she/her)

Sr. DX Engineer, Galileo.ai

Confidential | © 2025 Galileo Technologies, Inc.

2025 | 1

2 of 21

Generating Laughter�Evaluating the Success of LLMs for Comedic Purposes

Erin Mikail Staples (she/her)

Sr. DX Engineer, Galileo.ai

Confidential | © 2025 Galileo Technologies, Inc.

2025 | 2

3 of 21

Erin Mikail Staples (she/her)

developer experience engineer,�open source enthusiast, + ai-enhanced stand-up comedian

erinmikailstaples.com�erin@galileo.ai

2025 | 3

4 of 21

2025 | 4

5 of 21

2025 | 5

6 of 21

The roadmap.

Nondeterminism Isn’t a Bug, It’s the Bit

Understanding Metrics

Building your metrics stack

Business Case Examples

Startup Simulator 3000

2025 | 6

7 of 21

Nondeterminism Isn’t a Bug, It’s the Bit

2025 | 7

8 of 21

LLMs are unpredictable by design

But how do you test something you don’t expect to repeat?

2025 | 8

9 of 21

Keeping One Foot

on the Ground

2025 | 9

10 of 21

Standard AI Evaluation Metrics

✅ Accuracy: Is it truthful?

🔢 Token Usage: How many tokens did it burn?

🕓 Latency: How long did it take?

🛟 Safety: Did the model create harm?

2025 | 10

11 of 21

Serious Mode

Silly Mode

Pitch Decks

News API

Serious Formatter

Hacker News

Satire

Silly Formatter

Confidential | © 2025 Galileo Technologies, Inc.

2025 | 11

12 of 21

Traditional �AI Eval �Metrics

Agentic Experience

Metrics

✨�What really

matters

2025 | 12

13 of 21

Traditional �AI Eval �Metrics

Agentic Experience

Metrics

✨�What really

matters

Custom Metrics (involve SME)

2025 | 13

14 of 21

Beyond Comedy:

Custom metrics for

any industry.

2025 | 14

15 of 21

What should I measure?

Start with key metrics - Focus on metrics most relevant to your use case.

2025 | 15

16 of 21

What should I measure?

Establish baselines - Understand your current performance before making changes

2025 | 16

17 of 21

What should I measure?

Track trends over time - Monitor how metrics change as you iterate on your system

2025 | 17

18 of 21

What should I measure?

Combine multiple metrics - Look at related metrics together for a more complete picture

2025 | 18

19 of 21

What should I measure?

Set thresholds - Define acceptable ranges for critical metrics

2025 | 19

20 of 21

What should I measure?

Start with key metrics - Focus on metrics most relevant to your use case

Establish baselines - Understand your current performance before making changes

Track trends over time - Monitor how metrics change as you iterate on your system

Combine multiple metrics - Look at related metrics together for a more complete picture

Set thresholds - Define acceptable ranges for critical metrics

2025 | 20

21 of 21

Keep exploring!

Erin Mikail Staples (she/her)

erin@galileo.ai

https://git.new/Data-AI-Summit

Confidential | © 2025 Galileo Technologies, Inc.

2025 | 21