Ethical Machine Learning
Identifying and Removing Unfairness Caused by Biased Data
Dr. Pablo Rivas
Assistant Professor of Computer Science
School of Computer Science and Mathematics
http://www.rivas.ai
Pablo.Rivas@Marist.edu
Assumptions: Preprocess
If there are clearly sensitive attributes in the data, they must be removed
Assumptions: Unstability
If the training algorithm is known to be unstable, do something about it
Individual Fairness: Individuals who are similar should be treated similarly
Group Fairness: Different groups should be treated the same on average
In machine learning, do we care more about group fairness or individual fairness?
[R. Courtland, 2018]
Industry Problem
Researchers are pushing ahead on strategies for detecting bias in algorithms that haven’t been opened up for public scrutiny.
Rachel Courtland
Firms might be unwilling to discuss how they are working to address fairness, because it would mean admitting that there was a problem in the first place.
Journalist @ Nature
The solution
We need to care for both:
How it works in reality
Example: Binary Classification of People at Risk of Tax Cheating
Decision Maker
Citizen
Society
Of those classified as high-risk, how many will cheat?�Positive Predictive Value
What are the chances that I will be incorrectly classified as high-risk?�False Positive Rate
Is the selected training set demographically balanced?�Probability of Selection
PPV
FPR
Diversity
Is there a way to fix this?
Absolutely, it’s easy.
Let's talk about it!
Probability of Selection
Train ML Model
Demographic Parity
A group fairness problem in which we are concerned of demographic diversity
[Lidia T. Liu 2018]
Maximize for
Demographic Parity
We have to be careful about this types of impositions on classification systems because we cannot foresee the impact a fairness criterion would have if enforced.
Lidia T. Liu
However, if such an accurate model is available, we must use it. And it is likely that there are more direct ways to optimize for positive long-term outcomes than just Demographic Parity.
RISElab @ UC Berkeley
PPV: Positive Predictive Value
and
NPV: Negative Predictive Value
Train ML Model
Predictive Parity
A group fairness problem in which we desire for all groups to have equal PPV and NPV
[Alexandra Chouldechova 2016]
Maximize for
Predictive Parity
Predictive parity amounts to requiring that the positive PPV of the classifier be the same across all groups. However, predictive parity and demographic parity are not the same.
Alexandra Chouldechova
Demographic parity scores can fail to satisfy predictive parity because the relationship between data distributions can differ across groups in ways that result in PPV or NPV imbalance.
Assistant Professor @ Carnegie Mellon University
False Positive Rates
Train ML Model
Error Rate Balance
A group fairness problem in which we are concerned that classification systems make mistakes in the same proportion regardless of the group to which they belong
[Sorelle Friedler et.al. 2019]
Maximize for
Error Rate Balance
The aim is to not only to achieve equal positive predictive values values across different groups, but further equalize the odds of achieving equal errors across different groups.
Sorelle A. Friedler
In general, we can think of fairness in terms of group-conditioned error rates. But once we have achieved this we have to consistently consider the following moving forward, from a corrective approach, to a prescriptive approach
Assistant Professor @ Haverford College
Demo
Google Colaboratory
https://colab.research.google.com/drive/10hnqIxPD0g2Dn8RrLf1fBd-v_OdHs8Un
Ethical Machine Learning
Exponential Interest Due to Wide Accessibility of ML and Lack of Accountability
ML Fairness?
Face Recognition
Credit Worthiness
AI Ethicists
Ethical ML
We cared more about ability to train on big data
Face recognition problems were evident
BIased ML models were exposed publicly
ML Fairness and Algorithmic Fairness introduced to the public
All major ML/AI conf. have an ethics track
2010 | 2011 | 2012 | 2013 | 2014 | 2015 | 2016 | 2017 | 2018 | 2019 | 2020 | 2021 |
Thanks for listening
Questions?
Confusion Matrix