1 of 36

A Design Laboratory:

Language and Language Models

Karina Nguyen

Twitter Email

Jacobs Institute for Design Innovation at UC Berkeley

October 24, 2023

2 of 36

A Design Laboratory:

Language and Language Models

  1. Language & cultures of writing
  2. Language as a high-end product (The NYT)
    1. Elections 2020
    2. Journalism tooling
  3. Language as a committed way of showing love to users (Dropbox, Square)
  4. Language as the most compressed artifact of human knowledge (Anthropic)
    1. Interface innovations / paradigm shift
    2. Finding right form factors for novel capabilities
    3. Delightful micro-experiences
    4. Research storytelling

Jacobs Institute for Design Innovation

October 24, 2023

3 of 36

Language and cultures of writing.

4 of 36

Language and cultures of writing.

“Language is one of the most interesting laboratories to study things.” – Dario Amodei

5 of 36

Language and cultures of writing.

Language evolves along with culture. New words are created to describe new concepts, technologies, social changes etc.

Language can shape and influence the way people in a culture think and perceive the world, a concept known as linguistic relativity or the Sapir-Whorf hypothesis. For instance, the presence or absence of certain vocabulary can emphasize or downplay certain concepts or feelings, thereby molding perceptions and behaviors.

  • Programming languages
  • Natural language
  • Language models

6 of 36

Language as a high end product.

Most effective way to get people to read more of storylines coverage is by adding a layer of context to our product and coverage — across major surfaces — that connects people to the stories and information they need.

Elections 2020

7 of 36

8 of 36

9 of 36

Language as a high end product.

The goal is to increase the number of elections readers who read 2+ stories by 5% overall. This would lead to about 100,000 more readers a day reading a second elections article.

Context Blocks

10 of 36

Language as a high end product.

11 of 36

Language as a high end product.

Journalist Workflow

12 of 36

Language as a high end product.

iOS Widgets for Breaking News

13 of 36

Language as a high end product.

Engineering Catalog

14 of 36

Language as a committed way of showing love to users

15 of 36

Language as a committed way of showing love to users

Dropbox

16 of 36

Language as a committed way of showing love to users

Square – Virtual Terminal

17 of 36

Language as a committed way of showing love to users

Square

18 of 36

Language as the most compressed artifact of human knowledge

19 of 36

100K Context

Language as the most compressed artifact of human knowledge

20 of 36

Language as the most compressed artifact of human knowledge

21 of 36

Claude.ai – File Uploads

Language as the most compressed artifact of human knowledge

22 of 36

Language as the most compressed artifact of human knowledge

23 of 36

Claude.ai – Auto-generating Titles

Language as the most compressed artifact of human knowledge

24 of 36

Language as the most compressed artifact of human knowledge

25 of 36

Model Safety Card – Claude 2

Case Studies

Audience: policymakers, journalists, developers, academics

Main story: Claude 2 was an incremental improvement / evolution from previous model that steered towards what people wanted / user needs

  • This shaped some of our design decisions

Details

  • Consistent color palette, typography
  • Structure == storyline

26 of 36

Claude Instant 1.2

Case Studies

Audience: developers, companies, journalists

Main story: cheaper, faster, better and worse than Claude 2 (cost-effective)

Note: 80/20, no model card but lots of qual. evals

27 of 36

Measuring Subjective Global Opinions in LLMs

Case Studies

Audience: academics, policymakers

Main story:

  • New eval
  • Claude is more liberal

Details

Interactive data vis: llmglobalvalues.anthropic.com

28 of 36

Methods

Time

Heatmap

Legend

29 of 36

Process

Measuring Subjective Global Opinions in LLMs

  1. Sketching / Storyboarding
  2. Scoping
    1. Constraints (time)
    2. Main interaction points
    3. Map scale?
  3. Building out the MVP
    • React app + d3.js tooling
  4. Refining
    • We had multiple models selector, removed to be consistent with the paper
    • Adding more details (e.g. explain methodology, clarifying language)

30 of 36

Measuring Subjective Global Opinions in LLMs

31 of 36

Discovering LM’s Behavior Using Model-Written Evals

How I use Claude for research storytelling

Audience: academics, developers, policymakers

Main story: Model-written evals are as diverse and high quality as human manual evals but cheaper and more scalable. We generated ~150 evals and find that LMs are “sycophantic” due to RLHF and as the model size increases.

Details

  • Interactive data vis: evals.anthropic.com

32 of 36

Evaluations

Main story: Diversity within each dataset

Details

  • Scrollable sidebar to look at each category of eval
    • Horizontal scale + zoom in
  • More details on hover
  • Auto-generated groupings
  • Color choices
  • It was important to not overwhelm users

How I use Claude for research storytelling

33 of 36

How I use Claude for research storytelling

34 of 36

Language as the most compressed artifact of human knowledge

35 of 36

If Stripe is a monstrously successful business but what we make isn't beautiful and Stripe doesn't embody a culture of incredibly exacting craftsmanship, I'll be much less happy.

My intuition is that more of Stripe's success than one would think is downstream of the fact that people like beautiful things. Because what does a beautiful thing tell you?

Well, it tells you the person who made it really cared."

Patrick Collison

36 of 36

Thanks! Questions?

Twitter: @karinanguyen_

Email: karinanguyen28@gmail.com

Jacobs Institute for Design Innovation UC Berkeley

October 24, 2023