1 of 41

Uncertainty intervals for ranking

Yuval Benjamini

Hebrew University of Jerusalem, Israel

ISSI March 2026

1

2 of 41

Joint work with …

2

Bitya Neuhof

Hebrew University

Yoav Benjamini

Tel Aviv University

3 of 41

Problem overview

  •  

3

4 of 41

Machine learning examples I: Ranking Feature Importance

  •  

4

Explaining by Removing, Covett et al 2020

5 of 41

Machine learning examples I: Ranking Feature Importance

  • FI is usually reported by rank�”in the top 5”�“most important features”�“a graph of the top 10”.

  • Because FI values sensitive to �ML algorithm, �modification type, �loss / prediction.�Rank interpretation is robust.

5

“Using Shapley additive explanation values to estimate relative variable importance, we show that policy indicators—climate policy, scenario archetypes and technology policy—are ranked among the five most influential variables for the full ensemble of scenarios.”

Jaxa-Rozen, M., & Trutnevyte, E. (2021). Nature Climate Change

6 of 41

Machine learning examples I: Ranking Feature Importance

6

Working example based on �Capital Bike Share Dataset �for predicting total rental counts.

Uncertainty reported in FI value, when reported

Even though the ranking can also have errors / uncertainty

7 of 41

Machine learning examples II: Model leaderboards

  • Public model leaderboards map �architecture and company tradeoffs. �Usually reported in rank and score �(for top models)

7

8 of 41

Machine learning examples II: Model leaderboards

  • Public model leaderboards map �architecture and company tradeoffs. �Usually reported in rank and score �(for top models)

  • Each leaderboard is based on multiple tasks�with scale unknown and unimportant�also typically reported as ranks

8

9 of 41

Goal: Uncertainty in rankings

9

In both cases, rankings are reported:

  • Explicitly: “Top model”, “In top 5”
  • Implicitly: Sort the graph; Show only top k

But uncertainty shown in score-space.

We want to add uncertainty estimates in rank-space.

Score Uncertainty

Rank Uncertainty

10 of 41

Some distributional properties to note

10

Varies by ML algorithm, FI algorithm, and data distribution, but generally:

  • FI Values / Model scores can have relatively long tails
  • Deviation may increase with mean
  • Non-negligible correlation in dataset

11 of 41

Overview for rest of the talk

11

  • Need for uncertainty estimates ranks in ML
  • Pairwise testing approach to estimate rank confidence intervals
  • FWER based marginal and simultaneous intervals�[Based on Neuhof and Benjamini, 2024]
  • Leaky confidence intervals using FDR-controlling procedures
  • Leaderboard-level rank prediction intervals�

12 of 41

Definitions: True rank sets

12

 

13 of 41

A (marginal) rank confidence interval for unit j

  •  

13

14 of 41

Rank CIs from pairwise tests - a general method

  •  

14

 

 

15 of 41

Family wise error for marginal rank CIs

15

 

16 of 41

Rank CIs using FWER controlling methods

16

 

 

17 of 41

Comment: Naive bootstrap does not produce correct rank CIs

  •  

17

18 of 41

Simultaneous rank confidence interval set

  •  

18

19 of 41

Simultaneous rank confidence interval set

  •  

19

  • Is tuned for independent tests
  • BUT:�Adaptivity important �Accounts for dependent Xs
  • Accounts for dependent tests
  • BUT:�Not adaptive for rejections�Tuned for independent Xs

20 of 41

Simulations

20

Simulations

Example

 

 

Best case is 0.82

21 of 41

Simulation studies: Normalized average CI length

21

As correlation Increases, naive Tukey method becomes less efficient

no correlation and large variance -> Tukey + Min-P are better than Holm

p = 30

1-⍺ = 0.9

Variances uneven

Equal correlations

 

22 of 41

Pointers to other methods / approaches

  • CSranks by Mogstad, Romano, Shaikh, and Wilhelm have a bootstrap interface�for R even for correlated data. Note that it assumes normality with known covariance.
  • Comparisons for simultaneous intervals �were to the nominal Tukey implemented in ICranks �(Al Mohamad et al).
  • We can try to limit orders that are tested:
    • Goldwasser and Hooker have a top-to-bottom procedure for FI (ignore correlations)
    • Can use sample splitting to choose promising orders (see Kim and Ramdas) �or otherwise limit comparison (see discussion in our paper).
  • Partition-based methods “partition” the space of orderings. �They are statistically more efficient, but hard computationally. �Often use partial-conjunction based statsitics (tuned for iid case). �(Al Mohamad, van Zwet, Solari, Goeman; Heller and Solari)

22

 

23 of 41

Methods form conservative intervals (even if you bootstrap).

Why are intervals too large:

    • Don’t understand pairwise structure and dependence correctly (outside of Tukey)
    • Don’t understand role of correlation correctly
    • A theoretical alpha/2 for sign determination that is almost always conservative
    • These can be corrected by proper bootstrapping

    • FWE inherently a conservative measure. Can we use FDR ?

    • We don’t need to identify the rejected hypotheses, only to count them

23

 

24 of 41

Overview

24

  • Examples of rank-intervals in ML
  • Pairwise testing approach to estimate rank confidence intervals
  • FWER based marginal and simultaneous intervals [Neuhof and Benjamini, 2024]
  • Rank intervals with FDR-like control
  • Leaderboard-level rank prediction intervals

25 of 41

BH procedure for pairwise tests

  •  

25

 

26 of 41

Using BH procedures for pairwise tests

  •  

26

 

27 of 41

Bound severity and bound violations

Severity of bound is number of rejected hypotheses

� Lower (upper) bound violation is the size of the coverage error

Severity is equivalent to rejections in false-discovery rate.

Violation <= number of false rejections.

27

n

 

 

 

 

28 of 41

FDR-like methods: Violation rate controlling intervals

  •  

28

 

29 of 41

FDR-like methods: Violation rate controlling intervals

  •  

29

 

30 of 41

Simulations

Example

Simulations

 

 

We see pronounced improvement in power in the BH procedure

Intervals are still too long, probably…

31 of 41

Additional results

1. Using adaptive algorithms to improve bounds when few-nulls. � Among adaptive procedures, BKY has tight control of FDR for many-to-one comparisons.� This improves rank intervals for very high / very low ranked methods.

2. Selection:� When estimating intervals for all units, � BH-based intervals offer average control on VR rates. � � A. filtering of low-performing tasks (currently uses sample splitting).

B. adapting guarantees for average control over selected � when selection based on the intervals.

31

 

32 of 41

Overview

32

  • Need for intervals for ranks in ML
  • Pairwise testing approach to estimate rank confidence intervals
  • FWER based marginal and simultaneous intervals�[Based on Neuhof and Benjamini, 2024]
  • Leaky confidence intervals using FDR-controlling procedures
  • Leaderboard-level rank prediction intervals

33 of 41

Model leaderboards

  • We’ll use data from TabArena [Erickson et al.]
  • Benchmark system for prediction with tabular datasets.
  • Overall comparison using Elo ratings

33

  • 51 datasets
  • 50 models

34 of 41

Intervals for leaderboards

34

 

 

Leaderboard

Ranks CIs

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

Observed

Ranks

Observed Performance

35 of 41

Intervals for leaderboards

35

 

In practice tasks can be very diverse.

Model quality changes between tasks. �Should we strive for a single rank per?

All Model rank CIs for task

Tasks

Model CI

36 of 41

Intervals for leaderboards

36

 

 

Leaderboard

Ranks Interval

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

Task

Ranks CIs

Observed

Ranks

Observed Performance

37 of 41

Rank Prediction Intervals for leaderboards

37

 

 

38 of 41

Estimating prediction intervals

38

 

 

(*) Setting� avoids trivial quantiles

* See de Paula’s talk in April 9 on �Conformal Sets for intervals

39 of 41

Example on subset of TabArena

39

 

 

Based on task rank VR intervals

Based on task rank CI

Leaderboard rank-intervals

40 of 41

Thank You !

More Information:

  • Confident Feature Rankings in PMLR 2024 with python code�with more discussion about feature importance.
  • Our paper on rank-intervals for tasks, prediction intervals for leaderboard� is under review. Look for arxiv version this week. �If you’re interested in leaderboard evaluation, please reach out !
  • Work on BH should be out in near future.

Yuval.Benjamini@mail.huji.ac.il

40

 

41 of 41

Summary till now

41

  • We identify methods to build simultaneous CI for rankings
  • The methods efficiency depends on the testing:
    • Tukey-like tests efficient for iid with similar variances
    • Paired t-tests with Holm efficient for stronger correlations
    • Resampling based methods adaptive (but slow)
  • All methods are very conservative.

Note: In practice, for FI data Gaussian assumption not always met. � Can add robust tests and let min-p handle this.