1 of 40

SOTOPIA: INTERACTIVE EVALUATION FOR SOCIAL INTELLIGENCE IN LANGUAGE AGENTS

as presented by Tripp Whaley

2 of 40

Overview

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

3 of 40

Motivation

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

4 of 40

Humans are social, LLMs should be too

  • Dynamic and goal-driven social intelligence in LLMs
  • Balance a goal with multifaceted social norms
  • Simulate multi-turn conversations
    • Evaluate adherence to social norms

5 of 40

What is an Agent?

Agent - a system augmenting an LLM with additional capabilities, often including memory, action planning, or tool usage

6 of 40

Proposal: SOTOPIA

  1. Realistic - situations and characters are believable
  2. Mixed utilities - account for explicit and implicit goals
  3. Open-Ended - allow for flexibility in number of turns in conversation

​

7 of 40

Proposal: SOTOPIA

  1. Realistic - situations and characters’ relationships are grounded in reality
  2. Mixed utilities - account for explicit and implicit goals
  3. Open-Ended - allow for flexibility in number of turns in conversation

​

8 of 40

Proposal: SOTOPIA

  1. Realistic - situations and characters’ relationships are grounded in reality
  2. Mixed utilities - account for explicit and implicit goals
  3. Open-Ended - allow for flexibility in number of turns in conversation

​

9 of 40

Proposal: SOTOPIA

  1. Realistic - situations and characters’ relationships are grounded in reality
  2. Mixed utilities - account for explicit and implicit goals
  3. Open-Ended - scalable and procedurally generatible

​

10 of 40

Proposed Solution

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

11 of 40

SOTOPIA: Task Generation

  1. Characters - 40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

​

Episode

12 of 40

SOTOPIA: Task Generation

  1. Characters �40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

​

13 of 40

SOTOPIA: Task Generation

  1. Characters �40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

​

  • Name
  • Gender
  • Age
  • Occupation
  • Pronouns
  • Secret information
  • Public information
  • Personality traits
  • Moral values
  • Decision-making style
  • Schwartz personal values

14 of 40

SOTOPIA: Task Generation

  1. Characters �40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

​

  • Name
  • Gender
  • Age
  • Occupation
  • Pronouns
  • Secret information
  • Public information
  • Personality traits - Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism
  • Moral values - Care, Fairness, Loyalty, Authority, Purity
  • Decision-making style - Directive, Analytical, Conceptual, Behavioral
  • Schwartz personal values - Self-direction, Simulation, Hedonism, Achievement, Power, Security, Conformity, Tradition, � Benevolence, Universalism

15 of 40

SOTOPIA: Task Generation

  1. Characters - 40 characters generated from GPT-4 with attributes…
  2. Relationships -�Sample 90 pairs of characters�Assign relationship from:�family, friend, romantic, acquaintance, stranger�� explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

​

16 of 40

SOTOPIA: Task Generation

  1. Characters - 40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios�Shared Information - context of episode events�Private Information - character goals / secrets - allow for flexibility in number of turns in conversation

​

17 of 40

SOTOPIA: Task Generation

  1. Characters - 40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios�Shared Information - context of episode events�Private Information - character goals / secrets - allow for flexibility in number of turns in conversation

​

Scenario: Negotiation ��Shared Information: One person is selling an antique chair for $100 on his patio. Another person is interested in this chair. ��Agent 1 Private Information: Your goal is to buy the chair for $80.

�Agent 2 Private Information: Your goal is to sell the chair for $90.��Conflicting goals of agents create scenario constraints

​

18 of 40

SOTOPIA: Task Generation

  1. Characters - 40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

​

Episode

19 of 40

SOTOPIA: Task Generation

  1. Characters - 40 characters generated from GPT-4 with attributes…
  2. Relationships - account for explicit and implicit goals
  3. Scenarios - allow for flexibility in number of turns in conversation

​

Episode��- 2 agents with unique social goals placed in a scenario (based on relationship)��- 20 turns (max)

20 of 40

SOTOPIA: Agent Actions

  1. Speak
  2. Non-verbal Communication
  3. Physical Actions
  4. Nothing
  5. Leave

21 of 40

SOTOPIA: Agent Actions

  1. Speak
  2. Non-verbal Communication
  3. Physical Actions
  4. Nothing
  5. Leave

22 of 40

Metrics

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

​

23 of 40

SOTOPIA-EVAL: Agent Evaluation

  • How to measure social intelligence?

​

​

  • Some interactions create positive influence, others negative

​

​

  • Not just evaluating completion of task, but the social awareness with which they did

24 of 40

SOTOPIA-EVAL: Agent Evaluation

Dimension

Abbrev.

Min Score

Max Score

Description

Negative Example with Seller Agent

Goal Completion

GOAL

0

10

Agent’s social goal success

Fails to sell chair

Believability

BEL

0

10

Realism, Alignment with character profile

Offers to reduce price 50%

Knowledge

KNO

0

10

Information gathering ability

Fails to find way to convince buyer

Secret

SEC

-10

0

Keep secret information secret

Tells minimum price to buyer

Relationship

REL

-5

5

Were relationships hurt in achieving goal?

​

Social Rules

SOC

-10

0

Follows social norms and laws

​

Financial and Material Benefits

FIN

-5

5

Creates short or long-term opportunity

​

25 of 40

SOTOPIA-EVAL: Agent Evaluation

Key takeaways:

  • Semantic meaning embedded in scoring
  • Minimum overall score (-30)
  • Maximum overall score (40)

26 of 40

Experiments

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

27 of 40

SOTOPIA: Experimental Setup

Answer two core questions:

  1. How well can GPT-4 evaluate agents’ social interaction compared to humans
  2. What are the differences among models and humans in goal-oriented social intelligence?

28 of 40

Results

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

29 of 40

  1. Can GPT-4 evaluate like humans?
  • Generally, GPT-4 evaluates within 1 SD of humans

30 of 40

  1. Can GPT-4 evaluate like humans?
  • Generally, GPT-4 evaluates within 1 SD of humans
  • GPT-4 is better for REL, FIN, GOAL with models
  • GPT-4 is better for GOAL with humans

Who is the agent?

31 of 40

2a. How do models compare?

  • Average score when partner is a different model
  • GPT-4 wins

32 of 40

2a. How do models compare?

  • Average score when partner is a different model
  • GPT-4 wins
  • Average score across all dimensions

33 of 40

2b. How do humans and models compare?

  • Define SOTOPIA-HARD
    • 20 most challenging scenarios
    • “Challenging” => larger gap between maximum and minimum rewards
    • Target model: GPT-4

​

​

  • Define experiment setup
    • 20 GPT-4 with Human partner agent
    • 20 Humans with GPT-4 partner agent
    • 20 Humans with Human partner agent

34 of 40

2b. How do humans and models compare?

Humans perform significantly better than GPT-4 in GOAL

35 of 40

2b. How do humans and models compare?

Humans perform significantly better than GPT-4 in GOAL

36 of 40

2b. How do humans and models compare?

Humans perform significantly better than GPT-4 in GOAL

Recall performance with other models on all tasks…

​

What do we think of these results?

37 of 40

Conclusions

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

​

38 of 40

SOTOPIA Conclusions

  • The authors presented SOTOPIA and SOTOPIA-EVAL
    • A realistic, interactive, scenario-based dialogue engine
    • Evaluation of LLMs vs LLMs, LLMs vs Humans
  • Interactive dialogue > static benchmarks
    • Llama2-70b-chat performance
  • Quality of conversation partner affects social intelligence scoring
    • GPT-4 helps MPT, and MPT hurts GPT-4
  • Multiple dimensions of evaluation necessary
    • All models revealed secrets or broke social norms
  • Models make creative solutions (sometimes)

39 of 40

Discussion

Overview

Motivation

Proposed Solution

Metrics

Experiments

Results

Conclusions

Discussion

​

40 of 40

Discussion Points

  1. Did the experimental setup feel clear?
  2. Thoughts on dimensions of SOTOPIA-EVAL?
  3. Why only evaluate humans on the hardest scenarios?
  4. Why did llama2-70B-chat perform worse than expected?