1 of 21

Ethical Machine Learning

Identifying and Removing Unfairness Caused by Biased Data

Dr. Pablo Rivas

Assistant Professor of Computer Science

School of Computer Science and Mathematics

http://www.rivas.ai

Pablo.Rivas@Marist.edu

2 of 21

Assumptions: Preprocess

If there are clearly sensitive attributes in the data, they must be removed

3 of 21

Assumptions: Unstability

If the training algorithm is known to be unstable, do something about it

4 of 21

5 of 21

Individual Fairness: Individuals who are similar should be treated similarly

6 of 21

Group Fairness: Different groups should be treated the same on average

7 of 21

In machine learning, do we care more about group fairness or individual fairness?

[R. Courtland, 2018]

8 of 21

Industry Problem

Researchers are pushing ahead on strategies for detecting bias in algorithms that haven’t been opened up for public scrutiny.

Rachel Courtland

Firms might be unwilling to discuss how they are working to address fairness, because it would mean admitting that there was a problem in the first place.

Journalist @ Nature

9 of 21

The solution

We need to care for both:

  • Groups
  • Individuals

10 of 21

How it works in reality

Example: Binary Classification of People at Risk of Tax Cheating

Decision Maker

Citizen

Society

Of those classified as high-risk, how many will cheat?�Positive Predictive Value

What are the chances that I will be incorrectly classified as high-risk?�False Positive Rate

Is the selected training set demographically balanced?�Probability of Selection

PPV

FPR

Diversity

11 of 21

Is there a way to fix this?

Absolutely, it’s easy.

Let's talk about it!

12 of 21

Probability of Selection

Train ML Model

Demographic Parity

A group fairness problem in which we are concerned of demographic diversity

[Lidia T. Liu 2018]

Maximize for

13 of 21

Demographic Parity

We have to be careful about this types of impositions on classification systems because we cannot foresee the impact a fairness criterion would have if enforced.

Lidia T. Liu

However, if such an accurate model is available, we must use it. And it is likely that there are more direct ways to optimize for positive long-term outcomes than just Demographic Parity.

RISElab @ UC Berkeley

14 of 21

PPV: Positive Predictive Value

and

NPV: Negative Predictive Value

Train ML Model

Predictive Parity

A group fairness problem in which we desire for all groups to have equal PPV and NPV

[Alexandra Chouldechova 2016]

Maximize for

15 of 21

Predictive Parity

Predictive parity amounts to requiring that the positive PPV of the classifier be the same across all groups. However, predictive parity and demographic parity are not the same.

Alexandra Chouldechova

Demographic parity scores can fail to satisfy predictive parity because the relationship between data distributions can differ across groups in ways that result in PPV or NPV imbalance.

Assistant Professor @ Carnegie Mellon University

16 of 21

False Positive Rates

Train ML Model

Error Rate Balance

A group fairness problem in which we are concerned that classification systems make mistakes in the same proportion regardless of the group to which they belong

[Sorelle Friedler et.al. 2019]

Maximize for

17 of 21

Error Rate Balance

The aim is to not only to achieve equal positive predictive values values across different groups, but further equalize the odds of achieving equal errors across different groups.

Sorelle A. Friedler

In general, we can think of fairness in terms of group-conditioned error rates. But once we have achieved this we have to consistently consider the following moving forward, from a corrective approach, to a prescriptive approach

  • Emphasize preprocessing requirements
  • Avoid proliferation of measures
  • Account for training instability

Assistant Professor @ Haverford College

18 of 21

Demo

Google Colaboratory

https://colab.research.google.com/drive/10hnqIxPD0g2Dn8RrLf1fBd-v_OdHs8Un

19 of 21

Ethical Machine Learning

Exponential Interest Due to Wide Accessibility of ML and Lack of Accountability

ML Fairness?

Face Recognition

Credit Worthiness

AI Ethicists

Ethical ML

We cared more about ability to train on big data

Face recognition problems were evident

BIased ML models were exposed publicly

ML Fairness and Algorithmic Fairness introduced to the public

All major ML/AI conf. have an ethics track

2010

2011

2012

2013

2014

2015

2016

2017

2018

2019

2020

2021

20 of 21

Thanks for listening

Questions?

21 of 21

Confusion Matrix