1 of 46

Bias, Fairness, and Ethics

Gabrielle Burlison, Ziang Ren, Tongshu Yang

CS 6471

1/19/2022

1

2 of 46

What is Bias?

  • A tendency, inclination, or prejudice toward or against something or someone
    • What is positive bias?
    • What is negative bias?
    • What is implicit bias?
      • Attitudes and/or stereotypes towards people that form without our conscious knowledge.
  • Implicit Association Test

2

3 of 46

3

4 of 46

Implicit Association Test

  • Some tests you may not be aware of:
    • Presidents ('Presidential Popularity' IAT). This IAT requires the ability to recognize photos of Joseph Biden and one or more previous presidents.
    • Weight ('Fat - Thin' IAT). This IAT requires the ability to distinguish faces of people who are obese and people who are thin. It often reveals an automatic preference for thin people relative to fat people.
    • Weapons ('Weapons - Harmless Objects' IAT). This IAT requires the ability to recognize White and Black faces, and images of weapons or harmless objects.

4

5 of 46

Why should we care about bias and ethics?

  • We all represent many different viewpoints on various issues
  • The programs that we create try to emulate human decisions, so if humans have implicit bias, so will the programs (via the training data).
    • Think about applications: resume screening, medical diagnosis, criminal justice

5

6 of 46

Outline

Papers:

  • Gender differences in recommendation letters for postdoctoral fellowships in geoscience, Dutt et al.
  • Detecting Emergent Intersectional Bias, Guo & Caliskan

Structure

  • Overview
  • Background & Related Works
  • Experiment & Results
  • Strengths and Weaknesses

6

7 of 46

Gender Differences in Recommendation Letters For Postdoctoral Fellowships In Geoscience

Kuheli Dutt, Danielle L. Pfa, Ariel F. Bernstein, Joseph S. Dillard, Caryn J. Block

7

8 of 46

Overview

  • Under-representation of women in STEM is well-known, but the geosciences, in particular, represent a men-dominated field
  • Since the content of a recommendation letter is personal and generally private, implicit biases of the recommender are more likely to surface
  • Regional differences in letter length and gender differences in letter tone

8

9 of 46

Backgrounds

Women are under-represented in STEM disciplines.

  • 41% of STEM doctoral degrees awarded, but only occupy 24% postdoctoral
  • 35% less likely to get a tenure-track position than men
  • Hold fewer than 10% of full professor positions in the geosciences

9

Female Applicant

Male Applicant

Total

Female Recommender

67

81

148

Male Recommender

295

781

1076

Total

362

862

1224

Table 1 Recommendation Letters by Gender

10 of 46

Discussion

What are some reasons and consequences of females being underrepresented in STEM fields?

10

11 of 46

Backgrounds

Why recommendation letter?

  • Overall perception of a candidate’s ‘fit’ for a position
  • First impression of the applicant
  • Offer personal information about the candidate
  • Biases of the writer are more likely to surface

11

12 of 46

Procedure

Variables to Consider

12

Letter Length

Long Letter: > 50 lines

Short Letter: <= 10 lines

Letter Tone

Excellent: Accomplishment

Good: Acknowledgement

Doubtful: Uncertainty

Applicant’s Gender

Recommender’s Gender

Recommender’s Region

13 of 46

Results

Letter Length & Region

  • Americas have 56.7%
  • Americas’ letter length is much longer than other areas in average
  • No statistical difference of letter length between female and male applicants

13

Region

N

Mean

s.d.

Min

Max

Africa and Middle East

46

304.76

238.96

98

1,074

Australia, Europe, and New Zealand

253

345.05

187.42

60

986

South Asia

110

274.56

127.64

52

745

East Asia and Pacific

121

319.64

133.92

101

858

The Americas

694

561.06

311.49

37

2,444

Total

1,224

457.16

286.36

37

2,444

Table 2: Mean Letter Length by Region

14 of 46

Results

Letter Tone & Gender

  • Female applicants were half likely to receive excellent versus good letter
  • No statistical differences between male and female recommender
  • Male and female recommender did not differ likelihood to applicants

14

Excellent

Good

Doubtful

Total

Female applicant

53 (15%)

302 (83%)

7 (2%)

362

Male applicant

203 (24%)

635 (73%)

24 (3%)

862

1,224

Table 2: Letter Tone by Applicant Gender

15 of 46

Strengths

  • Largest study of gender bias in recommendation letters in STEM
  • Taken international data and the context of letter into consideration
  • Letter tone as one variable for testing
  • Analyzes many different combinations of variables to uncover correlations
    • Letter length, applicant gender, recommender gender, letter tone, region

15

16 of 46

Weaknesses

  • Primarily just a statistical analysis
  • Lack of Control for:
    • applicants’ qualifications
    • student-teacher relationships
  • More samples are needed to analyze the correlations of “doubtful” letters
  • Detailed linguistic analysis is needed

16

17 of 46

Key Concepts

  • Female applicants are much less likely to receive excellent recommendation letter compare to males
  • Recommender gender is not significant
  • Long letters are more likely to be excellent
  • Letter tone equivalently distributed across all regions

17

18 of 46

Relation to CSS

  • When doing social science research analysis, we need to pay attention to whether location would affect the result. Additionally, we could gather global data to study the effect of regional differences.

  • When doing language processing, the tone and context of the sentence could also provide useful information about the speaker which could be taken into consideration.

18

19 of 46

Connections Between Both Papers

19

20 of 46

Drawing Connections

  • Recommendation letters may shed some light on how bias gets integrated into the computer program
    • Programs reflect human biases
  • Gender differences… mentions that there is an “implicit gender bias” framing the issue, yet they have no way to quantitatively measure its impacts or effects
    • Therefore, bring in paper 2

20

21 of 46

Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases

Guo, Wei, and Aylin Caliskan

21

22 of 46

Drawing Connections

  • Both have room for future work and/or shortcomings
    • Binary gender
    • Continuous results
    • Limitations with respect to funding, energy, carbon footprint, data collection, etc.

22

23 of 46

Discussion

“The bias problem is present throughout machine learning approaches. As long as model is trained on biased information, the model will adopt those deviations in its learned parameters. This is the expected behavior of a machine learning algorithm: to mimic reality by identifying patterns in data.”

Let’s hear from both sides…

23

24 of 46

Overview

  • Implicit human biases are reflected in the statistical patterns in languages
  • Many static word embeddings (SWEs) are trained on natural language corpora
  • Off-the-shelf natural language models use static word embeddings trained on natural language corpora
  • Such natural language models are likely to contain bias that were present in
    • the corpora on which the models were trained
    • the corpora on which the SWEs used by the models are trained

24

25 of 46

Overview

  • Describe association with bias
    • (extended) Word Embedding Association Test (WEAT)
    • (extended) Word Embedding Factual Association Test (WEFAT)
  • Identifying different types of biases (from SWEs)
    • Intersectional Bias Detection (IBD)
    • Emergent Intersectional Bias Detection (EIBD)
  • Quantifying and measuring biases in the trained natural language models
    • Contextualized Embedding Association Test (CEAT)

25

26 of 46

Related Concepts

  • Word Embedding Association Test (WEAT)

Null hypothesis: no difference between the two sets of target words in terms of relative similarity to the attribute words

Effect Size (strength of the difference)

26

27 of 46

Related Concepts

  • Word Embedding Factual Association Test (WEFAT)
    • normalized difference between cosine similarities

27

28 of 46

Related Concepts

  • Intersectional Bias
    • Biases that are caused by multiple factors of advantage and disadvantage
    • Gender, race, sexuality, disability, etc.
  • Emergent Intersectional Bias
    • unique biases that arises when a person belongs to two or more disadvantage groups at the same time

Discuss:

  • What may be an intersectional bias based on these disadvantage groups
  • What may be an emergent intersectional bias?

28

29 of 46

Related Concepts

  • Factors of disadvantage
    • gender, race, sexuality, disability, weight, etc.
  • Intersectional Bias
    • Being African American and Female at the same time (DeGraffenreid v. General Motors 1976)
  • Emergent Intersectional Bias
    • Hair Weaves are typically associated with African American females but not males (Guo & Caliskan)

29

30 of 46

Related Concepts

  • Static Word Embeddings (SWEs)
    • Dense numerical vector representation of words
    • Example: Word2Vec, GloVe
  • Contextualized Word Embeddings (CWEs)
    • Dynamic word representations generated by natural language models that adapts to the context

30

31 of 46

Related Concepts

31

ELMo

2 layer Bi-LSTM

Billion Word Benchmark

93.6 million parameters

Integrates hidden states in all layers

BERT

Bidirectional Transformer encoder + Masked Language Model & Next Sentence Prediction

BookCorpus & English Wikipedia

12 layers (BERT-small-case) 110 million parameters

Uses hidden states in the top layer

GPT

12-layer Transformer Decoder + Unidirectional Language Model

BookCorpus

110 million parameters

Uses hidden states in the top layer

GPT-2

Transformer Decoder + Unidirectional Language Model

WebText

12 layers (GPT-2-small) 117 million parameters

Uses hidden states in the top layer

32 of 46

Methods - IBD

  • Intersectional Bias Detection
    • using a statistic analogous to WEFAT

  • A threshold t
  • Greater than threshold means w is associated with A

32

(Association Score)

33 of 46

Methods - IBD

For a group of intersections by two categories C1, C2 with m and n subcategories (M×N intersections in total), detect bias associated with group C11

  • Compute the association scores for words in each pair of categories (C11, Cij)
  • Compare the scores against a threshold (hyperparameter)
  • Collect the words associated with C11 in each pair
    • These words came from a collection of intersectional biased words for each pair and some random words with no bias

This is essentially a one-vs-all classifier model!

33

34 of 46

Methods - EIBD

For a group of intersections by two categories C1n, Cm1 with m and n subcategories (M×N intersections in total), detect emergent bias associated with group C11

  • Using IBD, compute a list of biased words associated with group C11
  • For each of the m subcategories Sin in C1n and n subcategories Smj in Cm1, Compute association score for the pair (S1n, Sin) and (Sm1, Smj)
  • Remove words that has a high association score in each pair

Removing words that are too strongly associated with a single constituent subcategory

34

35 of 46

Methods - EIBD

For a group of intersections by two categories C1n, Cm1 with m and n subcategories (M×N intersections in total), detect emergent bias associated with group C11

  • Using IBD, compute a list of biased words associated with group C11
  • For each of the m subcategories Sin in C1n and n subcategories Smj in Cm1, Compute association score for the pair (S1n, Sin) and (Sm1, Smj)
  • Remove words that has a high association score in each pair

Removing words that are too strongly associated with a single constituent subcategory

35

36 of 46

Methods - CEAT

Constructs a population of CWEs of all stimuli that we are interested in.

Given ns input sentences, calculate CWEs using a natural language model for each stimulus

For each stimulus

  • Sample random combinations of CWEs N times

For each sample

  • Sample without replacement if stimulus appeared at least N times
  • Sample with replacement otherwise

Recursion!

  • Word Embedding Association Test (WEAT)

Null hypothesis: no difference between the two sets of target words in terms of relative similarity to the attribute words

36

37 of 46

Methods - CEAT

For each sample

  • calculate the effect size, variance
  • approximate the distribution of effect sizes using Combined Effect Size

Recursion!

  • Word Embedding Association Test (WEAT)

Null hypothesis: no difference between the two sets of target words in terms of relative similarity to the attribute words

37

38 of 46

Evaluation - IBD & EIBD

  • 98 Attributes
  • 2 gender groups (female, males)
  • 3 racial groups (African, European, Mexican Americans)
  • Random words from WEAT not associated with any group

38

39 of 46

Evaluation - IBD & EIBD

Recursion!

Selecting a threshold: Highest TPR/FPR ratio

  • Intersectional Bias Detection
    • using a statistic analogous to WEFAT

  • A threshold t
  • Greater than threshold means w is associated with A

39

(Association Score)

40 of 46

Evaluation - IBD & EIBD

40

41 of 46

Evaluation - IBD & EIBD

41

 

IBD

Random (IBD)

EIBD

Random (EIBD)

African American Females

81.60%

14.30%

84.70%

9.20%

Mexican American Females

82.70%

13.30%

65.30%

6.10%

42 of 46

Evaluation - CEAT

  • CEAT detects stronger intersectional bias than singular biases
  • Bias GPT-2 < GPT < BERT < ELMo
  • Higher contextualization, less bias

Recursion!

42

ELMo

2 layer Bi-LSTM

93.6 million parameters

BERT

Bidirectional Transformer encoder + Masked Language Model & Next Sentence Prediction

12 layers (BERT-small-case) 110 million parameters

GPT

12-layer Transformer Decoder + Unidirectional Language Model

110 million parameters

GPT-2

Transformer Decoder + Unidirectional Language Model

12 layers (GPT-2-small) 117 million parameters

43 of 46

Strengths

  • IBD and EIBD provide a concrete way to quantitatively measure social bias in SWE
  • CEAT quantitatively measures the level of bias in natural language models
  • IBD is extendable
    • Although the paper primarily focused on binary gender and race, IBD can be easily extended to other categories including sexuality, age, disability status

43

44 of 46

Weaknesses

  • Study is based on Reddit corpus and comes with bias
  • Categorical representations
    • Binary gender
    • Multiple races and ethnicities
  • Extending to new categories require human annotated data
    • Lexicon induction partially mitigates this, but no method currently exists
  • Does not scale very well to open-ended bias detection
  • Does not provide causal explanation to bias in SWEs, CWEs and natural language models

44

45 of 46

Reference

Guo, Wei, and Aylin Caliskan. "Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases." In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pp. 122-133. 2021.

Dutt, Kuheli, Danielle L. Pfaff, Ariel F. Bernstein, Joseph S. Dillard, and Caryn J. Block. "Gender differences in recommendation letters for postdoctoral fellowships in geoscience." Nature Geoscience 9, no. 11 (2016): 805-808.

45

46 of 46

Thank you for listening!

46