Bias, Fairness, and Ethics
Gabrielle Burlison, Ziang Ren, Tongshu Yang
CS 6471
1/19/2022
1
What is Bias?
2
3
Implicit Association Test
4
Why should we care about bias and ethics?
5
Outline
Papers:
Structure
6
Gender Differences in Recommendation Letters For Postdoctoral Fellowships In Geoscience
Kuheli Dutt, Danielle L. Pfa, Ariel F. Bernstein, Joseph S. Dillard, Caryn J. Block
7
Overview
8
Backgrounds
Women are under-represented in STEM disciplines.
9
| Female Applicant | Male Applicant | Total |
Female Recommender | 67 | 81 | 148 |
Male Recommender | 295 | 781 | 1076 |
Total | 362 | 862 | 1224 |
Table 1 Recommendation Letters by Gender
Discussion
What are some reasons and consequences of females being underrepresented in STEM fields?
10
Backgrounds
Why recommendation letter?
11
Procedure
Variables to Consider
12
Letter Length
Long Letter: > 50 lines
Short Letter: <= 10 lines
Letter Tone
Excellent: Accomplishment
Good: Acknowledgement
Doubtful: Uncertainty
Applicant’s Gender
Recommender’s Gender
Recommender’s Region
Results
Letter Length & Region
13
Region | N | Mean | s.d. | Min | Max |
Africa and Middle East | 46 | 304.76 | 238.96 | 98 | 1,074 |
Australia, Europe, and New Zealand | 253 | 345.05 | 187.42 | 60 | 986 |
South Asia | 110 | 274.56 | 127.64 | 52 | 745 |
East Asia and Pacific | 121 | 319.64 | 133.92 | 101 | 858 |
The Americas | 694 | 561.06 | 311.49 | 37 | 2,444 |
Total | 1,224 | 457.16 | 286.36 | 37 | 2,444 |
Table 2: Mean Letter Length by Region
Results
Letter Tone & Gender
14
| Excellent | Good | Doubtful | Total |
Female applicant | 53 (15%) | 302 (83%) | 7 (2%) | 362 |
Male applicant | 203 (24%) | 635 (73%) | 24 (3%) | 862 |
| | | | 1,224 |
Table 2: Letter Tone by Applicant Gender
Strengths
15
Weaknesses
16
Key Concepts
17
Relation to CSS
18
Connections Between Both Papers
19
Drawing Connections
20
Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases
Guo, Wei, and Aylin Caliskan
21
Drawing Connections
22
Discussion
“The bias problem is present throughout machine learning approaches. As long as model is trained on biased information, the model will adopt those deviations in its learned parameters. This is the expected behavior of a machine learning algorithm: to mimic reality by identifying patterns in data.”
Let’s hear from both sides…
23
Overview
24
Overview
25
Related Concepts
Null hypothesis: no difference between the two sets of target words in terms of relative similarity to the attribute words
Effect Size (strength of the difference)
26
Related Concepts
27
Related Concepts
Discuss:
28
Related Concepts
29
Related Concepts
30
Related Concepts
31
ELMo | 2 layer Bi-LSTM | Billion Word Benchmark | 93.6 million parameters | Integrates hidden states in all layers |
BERT | Bidirectional Transformer encoder + Masked Language Model & Next Sentence Prediction | BookCorpus & English Wikipedia | 12 layers (BERT-small-case) 110 million parameters | Uses hidden states in the top layer |
GPT | 12-layer Transformer Decoder + Unidirectional Language Model | BookCorpus | 110 million parameters | Uses hidden states in the top layer |
GPT-2 | Transformer Decoder + Unidirectional Language Model | WebText | 12 layers (GPT-2-small) 117 million parameters | Uses hidden states in the top layer |
Methods - IBD
32
(Association Score)
Methods - IBD
For a group of intersections by two categories C1, C2 with m and n subcategories (M×N intersections in total), detect bias associated with group C11
This is essentially a one-vs-all classifier model!
33
Methods - EIBD
For a group of intersections by two categories C1n, Cm1 with m and n subcategories (M×N intersections in total), detect emergent bias associated with group C11
Removing words that are too strongly associated with a single constituent subcategory
34
Methods - EIBD
For a group of intersections by two categories C1n, Cm1 with m and n subcategories (M×N intersections in total), detect emergent bias associated with group C11
Removing words that are too strongly associated with a single constituent subcategory
35
Methods - CEAT
Constructs a population of CWEs of all stimuli that we are interested in.
Given ns input sentences, calculate CWEs using a natural language model for each stimulus
For each stimulus
For each sample
Recursion!
Null hypothesis: no difference between the two sets of target words in terms of relative similarity to the attribute words
36
Methods - CEAT
For each sample
Recursion!
Null hypothesis: no difference between the two sets of target words in terms of relative similarity to the attribute words
37
Evaluation - IBD & EIBD
38
Evaluation - IBD & EIBD
Recursion!
Selecting a threshold: Highest TPR/FPR ratio
39
(Association Score)
Evaluation - IBD & EIBD
40
Evaluation - IBD & EIBD
41
| IBD | Random (IBD) | EIBD | Random (EIBD) |
African American Females | 81.60% | 14.30% | 84.70% | 9.20% |
Mexican American Females | 82.70% | 13.30% | 65.30% | 6.10% |
Evaluation - CEAT
Recursion!
42
ELMo | 2 layer Bi-LSTM | 93.6 million parameters |
BERT | Bidirectional Transformer encoder + Masked Language Model & Next Sentence Prediction | 12 layers (BERT-small-case) 110 million parameters |
GPT | 12-layer Transformer Decoder + Unidirectional Language Model | 110 million parameters |
GPT-2 | Transformer Decoder + Unidirectional Language Model | 12 layers (GPT-2-small) 117 million parameters |
Strengths
43
Weaknesses
44
Reference
Guo, Wei, and Aylin Caliskan. "Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases." In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pp. 122-133. 2021.
Dutt, Kuheli, Danielle L. Pfaff, Ariel F. Bernstein, Joseph S. Dillard, and Caryn J. Block. "Gender differences in recommendation letters for postdoctoral fellowships in geoscience." Nature Geoscience 9, no. 11 (2016): 805-808.
45
Thank you for listening!
46