1 of 40

SOTOPIA: INTERACTIVE EVALUATION FOR SOCIAL INTELLIGENCE IN LANGUAGE AGENTS

as presented by Tripp Whaley

2 of 40

Overview

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

3 of 40

Motivation

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

4 of 40

Humans are social, LLMs should be too

  • Dynamic and goal-driven social intelligence in LLMs
  • Balance a goal with multifaceted social norms
  • Simulate multi-turn conversations
    • Evaluate adherence to social norms

5 of 40

What is an Agent?

Agent - a system augmenting an LLM with additional capabilities, often including memory, action planning, or tool usage

6 of 40

Proposal: SOTOPIA

  1. Realistic - situations and characters are believable
  2. Mixed utilities - account for explicit and implicit goals
  3. Open-Ended - allow for flexibility in number of turns in conversation

7 of 40

Proposal: SOTOPIA

  1. Realistic - situations and characters’ relationships are grounded in reality
  2. Mixed utilities - account for explicit and implicit goals
  3. Open-Ended - allow for flexibility in number of turns in conversation

8 of 40

Proposal: SOTOPIA

  1. Realistic - situations and characters’ relationships are grounded in reality
  2. Mixed utilities - account for explicit and implicit goals
  3. Open-Ended - allow for flexibility in number of turns in conversation

9 of 40

Proposal: SOTOPIA

  1. Realistic - situations and characters’ relationships are grounded in reality
  2. Mixed utilities - account for explicit and implicit goals
  3. Open-Ended - scalable and procedurally generatible

10 of 40

Proposed Solution

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

11 of 40

SOTOPIA: Task Generation

  1. Characters - 40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

Episode

12 of 40

SOTOPIA: Task Generation

  1. Characters40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

13 of 40

SOTOPIA: Task Generation

  1. Characters40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

  • Name
  • Gender
  • Age
  • Occupation
  • Pronouns
  • Secret information
  • Public information
  • Personality traits
  • Moral values
  • Decision-making style
  • Schwartz personal values

14 of 40

SOTOPIA: Task Generation

  1. Characters40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

  • Name
  • Gender
  • Age
  • Occupation
  • Pronouns
  • Secret information
  • Public information
  • Personality traits - Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism
  • Moral values - Care, Fairness, Loyalty, Authority, Purity
  • Decision-making style - Directive, Analytical, Conceptual, Behavioral
  • Schwartz personal values - Self-direction, Simulation, Hedonism, Achievement, Power, Security, Conformity, Tradition, � Benevolence, Universalism

15 of 40

SOTOPIA: Task Generation

  1. Characters - 40 characters generated from GPT-4 with attributes…
  2. Relationships -�Sample 90 pairs of characters�Assign relationship from:family, friend, romantic, acquaintance, stranger�� explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

16 of 40

SOTOPIA: Task Generation

  1. Characters - 40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. ScenariosShared Information - context of episode events�Private Information - character goals / secrets - allow for flexibility in number of turns in conversation

17 of 40

SOTOPIA: Task Generation

  1. Characters - 40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. ScenariosShared Information - context of episode events�Private Information - character goals / secrets - allow for flexibility in number of turns in conversation

Scenario: Negotiation ��Shared Information: One person is selling an antique chair for $100 on his patio. Another person is interested in this chair. ��Agent 1 Private Information: Your goal is to buy the chair for $80.

�Agent 2 Private Information: Your goal is to sell the chair for $90.��Conflicting goals of agents create scenario constraints

18 of 40

SOTOPIA: Task Generation

  1. Characters - 40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

Episode

19 of 40

SOTOPIA: Task Generation

  1. Characters - 40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

Episode�- 2 agents with unique social goals placed in a scenario (based on relationship)��- 20 turns (max)

20 of 40

SOTOPIA: Agent Actions

  1. Speak
  2. Non-verbal Communication
  3. Physical Actions
  4. Nothing
  5. Leave

21 of 40

SOTOPIA: Agent Actions

  1. Speak
  2. Non-verbal Communication
  3. Physical Actions
  4. Nothing
  5. Leave

22 of 40

Metrics

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

23 of 40

SOTOPIA-EVAL: Agent Evaluation

  • How to measure social intelligence?

  • Some interactions create positive influence, others negative

  • Not just evaluating completion of task, but the social awareness with which they did

24 of 40

SOTOPIA-EVAL: Agent Evaluation

Dimension

Abbrev.

Min Score

Max Score

Description

Negative Example with Seller Agent

Goal Completion

GOAL

0

10

Agent’s social goal success

Fails to sell chair

Believability

BEL

0

10

Realism, Alignment with character profile

Offers to reduce price 50%

Knowledge

KNO

0

10

Information gathering ability

Fails to find way to convince buyer

Secret

SEC

-10

0

Keep secret information secret

Tells minimum price to buyer

Relationship

REL

-5

5

Were relationships hurt in achieving goal?

Social Rules

SOC

-10

0

Follows social norms and laws

Financial and Material Benefits

FIN

-5

5

Creates short or long-term opportunity

25 of 40

SOTOPIA-EVAL: Agent Evaluation

Key takeaways:

  • Semantic meaning embedded in scoring
  • Minimum overall score (-30)
  • Maximum overall score (40)

26 of 40

Experiments

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

27 of 40

SOTOPIA: Experimental Setup

Answer two core questions:

  1. How well can GPT-4 evaluate agents’ social interaction compared to humans
  2. What are the differences among models and humans in goal-oriented social intelligence?

28 of 40

Results

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

29 of 40

  1. Can GPT-4 evaluate like humans?
  • Generally, GPT-4 evaluates within 1 SD of humans

30 of 40

  1. Can GPT-4 evaluate like humans?
  • Generally, GPT-4 evaluates within 1 SD of humans
  • GPT-4 is better for REL, FIN, GOAL with models
  • GPT-4 is better for GOAL with humans

Who is the agent?

31 of 40

2a. How do models compare?

  • Average score when partner is a different model
  • GPT-4 wins

32 of 40

2a. How do models compare?

  • Average score when partner is a different model
  • GPT-4 wins
  • Average score across all dimensions

33 of 40

2b. How do humans and models compare?

  • Define SOTOPIA-HARD
    • 20 most challenging scenarios
    • “Challenging” => larger gap between maximum and minimum rewards
    • Target model: GPT-4

  • Define experiment setup
    • 20 GPT-4 with Human partner agent
    • 20 Humans with GPT-4 partner agent
    • 20 Humans with Human partner agent

34 of 40

2b. How do humans and models compare?

Humans perform significantly better than GPT-4 in GOAL

35 of 40

2b. How do humans and models compare?

Humans perform significantly better than GPT-4 in GOAL

36 of 40

2b. How do humans and models compare?

Humans perform significantly better than GPT-4 in GOAL

Recall performance with other models on all tasks…

What do we think of these results?

37 of 40

Conclusions

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

38 of 40

SOTOPIA Conclusions

  • The authors presented SOTOPIA and SOTOPIA-EVAL
    • A realistic, interactive, scenario-based dialogue engine
    • Evaluation of LLMs vs LLMs, LLMs vs Humans
  • Interactive dialogue > static benchmarks
    • Llama2-70b-chat performance
  • Quality of conversation partner affects social intelligence scoring
    • GPT-4 helps MPT, and MPT hurts GPT-4
  • Multiple dimensions of evaluation necessary
    • All models revealed secrets or broke social norms
  • Models make creative solutions (sometimes)

39 of 40

Discussion

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

40 of 40

Discussion Points

  1. Did the experimental setup feel clear?
  2. Thoughts on dimensions of SOTOPIA-EVAL?
  3. Why only evaluate humans on the hardest scenarios?
  4. Why did llama2-70B-chat perform worse than expected?