1 of 27

Robust and Equitable�Uncertainty Estimation

Aaron Roth

Joint work with

Osbert Bastani, Varun Gupta, Chris Jung, Mallesh Pai, George Noarov, Ramya Ramalingam, Rakesh Vohra

2 of 27

3 of 27

Prediction Sets

  • We have many black box methods that can make predictions.
  • Which predictions should we feel confident in?
  • Need a way of quantifying uncertainty. One way: prediction sets.

Image from Angelopoulos, Bates, Malik, Jordan ICLR 2021

Goal: The prediction set should contain the true label with probability (e.g.) 95%.

4 of 27

Conformal Prediction [Shafer, Vovk]

  •  

 

 

 

 

 

 

 

5 of 27

What’s wrong with marginal guarantees?

 

How sure are you?

 

Hmmm…

6 of 27

What’s wrong with marginal guarantees?

 

 

7 of 27

Marginal Guarantees.

 

But I’m part of a demographic group representing less than 5% of the population…

For disjoint groups, could just calibrate separately on each group

[Romano, Foygel-Barber, Sabatti, Candes ‘20]

But…

8 of 27

Marginal Guarantees.

For people with egg allergies and no history of smoking, the 95% prediction interval is [e, f].

What about for people like me?

For women with a family history of diabetes the 95% prediction interval is [c,d]

For African Americans under the age of 50 the 95% prediction interval is [a,b]

What does this mean for me?

9 of 27

Group Conditional Validity

 

10 of 27

Calibration (Testing an Oracular Forecaster)

 

11 of 27

Prediction Set Multivalidity�Quantile Analogue of Multicalibration [Hebert-Johnson, Kim, Reingold, Rothblum ’18]

 

12 of 27

Group Conditional Guarantees�(Generalizing Related Algorithms for Mean Multicalibration: Squared Error -> Pinball Loss)

 

13 of 27

 

Multivalid Guarantees�(Generalizing Related Algorithms for Mean Multicalibration: Squared Error -> Pinball Loss)

14 of 27

15 of 27

Groupwise Coverage

  • Real data: An income prediction task derived from 2018 California Census data.
  • Groups defined by race and binary gender.

16 of 27

Groupwise Threshold Calibration

  • Real data: An income prediction task derived from 2018 California Census data.
  • Groups defined by race and binary gender.

17 of 27

Additional Results

  • All of this works in the online adversarial setting too!
  • Can characterize exactly which statistics are possible to multicalibrate with respect to and which are not…
    • Yes: Means, Quantiles, …
    • No: Variance, Conditional Value at Risk, …
  • Can jointly multicalibrate pairs of statistics if one statistic is conditionally elicitable, conditional on the other.
    • Mean and Variance Together
    • Quantile and Conditional Value at Risk together…

18 of 27

Thank You.

Moment Multicalibration for Uncertainty Estimation. Joint work with Chris Jung, Changhwa Lee, Mallesh Pai, Rakesh Vohra. In COLT 2021.

Online Multivalid Learning: Means, Moments, and Prediction Intervals

Joint work with Varun Gupta, Chris Jung, George Noarov, Mallesh Pai. In ITCS 2022

Practical, Adversarial, Multivalid Conformal Prediction

Joint work with Osbert Bastani, Varun Gupta, Chris Jung, George Noarov, Ramya Ramalingam. In NeurIPS 2022

Batch Multivalid Conformal Prediction. Joint work with Chris Jung, George Noarov, Ramya Ramalingam. In ICLR 2023.

The Scope of Multicalibration: Characterizing Multicalibration via Property Elicitation. Joint work with George Noarov. In ICML 2023.

Monograph under construction: Uncertain: Modern Topics in Uncertainty Quantification

http://uncertaintyclass.com

19 of 27

Unanticipated Distribution Shift

  • Task: Predicting income from Census data
  • Conformity Score derived from Quantile Regression
  • Split conformal prediction calibrated on 2018 California data
  • For both methods: First presented 35k random points from 2018 California data, then 20k from 2018 Pennsylvania data

20 of 27

Average Coverage:

Split Conformal: 0.915

Our Method: 0.906

Average Prediction Set Size:

Split Conformal: 2.329

Our Method: 2.1

Image from Angelopoulos et al. ICLR 2021

 

21 of 27

Competitive with split conformal prediction even on its own turf.

  • A synthetic linear regression problem. Data is i.i.d. Must simultaneously train a linear regression model and form prediction intervals.
  • Split conformal prediction must split data into a training and calibration set to maintain exchangability (80% train 20% test)
  • We can train on all data.

22 of 27

Competitive with split conformal prediction even on its own turf.

  •  

23 of 27

Advantages of Our Method�(Even in the i.i.d setting)

  •  

24 of 27

Means

(Multicalibration)

Quantiles

(Multivalidity)

25 of 27

Beyond Means and Quantiles

  •  

26 of 27

Beyond Means and Quantiles�(c.f. Nicolas Lambert)

  •  

27 of 27

Beyond Means and Quantiles

  •