Uncertainty intervals for ranking
Yuval Benjamini
Hebrew University of Jerusalem, Israel
ISSI March 2026
1
Joint work with …
2
Bitya Neuhof
Hebrew University
Yoav Benjamini
Tel Aviv University
Problem overview
3
Machine learning examples I: Ranking Feature Importance
4
Explaining by Removing, Covett et al 2020
Machine learning examples I: Ranking Feature Importance
5
“Using Shapley additive explanation values to estimate relative variable importance, we show that policy indicators—climate policy, scenario archetypes and technology policy—are ranked among the five most influential variables for the full ensemble of scenarios.”
Jaxa-Rozen, M., & Trutnevyte, E. (2021). Nature Climate Change
Machine learning examples I: Ranking Feature Importance
6
Working example based on �Capital Bike Share Dataset �for predicting total rental counts.
Uncertainty reported in FI value, when reported
Even though the ranking can also have errors / uncertainty
Machine learning examples II: Model leaderboards
7
Machine learning examples II: Model leaderboards
8
Goal: Uncertainty in rankings
9
In both cases, rankings are reported:
But uncertainty shown in score-space.
We want to add uncertainty estimates in rank-space.
Score Uncertainty
Rank Uncertainty
Some distributional properties to note
10
Varies by ML algorithm, FI algorithm, and data distribution, but generally:
Overview for rest of the talk
11
Definitions: True rank sets
12
A (marginal) rank confidence interval for unit j
13
Rank CIs from pairwise tests - a general method
14
Family wise error for marginal rank CIs
15
Rank CIs using FWER controlling methods
16
Comment: Naive bootstrap does not produce correct rank CIs
17
Simultaneous rank confidence interval set
18
Simultaneous rank confidence interval set
19
Simulations
20
Simulations
Example
Best case is 0.82
Simulation studies: Normalized average CI length
21
As correlation Increases, naive Tukey method becomes less efficient
no correlation and large variance -> Tukey + Min-P are better than Holm
p = 30
1-⍺ = 0.9
Variances uneven
Equal correlations
Pointers to other methods / approaches
22
Methods form conservative intervals (even if you bootstrap).
Why are intervals too large:
23
Overview
24
BH procedure for pairwise tests
25
Using BH procedures for pairwise tests
26
Bound severity and bound violations
Severity of bound is number of rejected hypotheses
� Lower (upper) bound violation is the size of the coverage error
Severity is equivalent to rejections in false-discovery rate.
Violation <= number of false rejections.
27
n
FDR-like methods: Violation rate controlling intervals
28
FDR-like methods: Violation rate controlling intervals
29
Simulations
Example
Simulations
We see pronounced improvement in power in the BH procedure
Intervals are still too long, probably…
Additional results
1. Using adaptive algorithms to improve bounds when few-nulls. � Among adaptive procedures, BKY has tight control of FDR for many-to-one comparisons.� This improves rank intervals for very high / very low ranked methods.
2. Selection:� When estimating intervals for all units, � BH-based intervals offer average control on VR rates. � � A. filtering of low-performing tasks (currently uses sample splitting).
B. adapting guarantees for average control over selected � when selection based on the intervals.
31
Overview
32
Model leaderboards
33
Figure source https://huggingface.co/spaces/TabArena/leaderboard
Intervals for leaderboards
34
Leaderboard
Ranks CIs
⋯
⋯
⋯
⋯
⋯
⋯
⋯
Observed
Ranks
Observed Performance
Intervals for leaderboards
35
In practice tasks can be very diverse.
Model quality changes between tasks. �Should we strive for a single rank per?
All Model rank CIs for task
Tasks
Model CI
Intervals for leaderboards
36
Leaderboard
Ranks Interval
⋯
⋯
⋯
⋯
⋯
⋯
⋯
⋯
⋯
⋯
Task
Ranks CIs
Observed
Ranks
Observed Performance
Rank Prediction Intervals for leaderboards
37
Estimating prediction intervals
38
(*) Setting� avoids trivial quantiles
* See de Paula’s talk in April 9 on �Conformal Sets for intervals
Example on subset of TabArena
39
Based on task rank VR intervals
Based on task rank CI
Leaderboard rank-intervals
Thank You !
More Information:
Yuval.Benjamini@mail.huji.ac.il
40
Summary till now
41
Note: In practice, for FI data Gaussian assumption not always met. � Can add robust tests and let min-p handle this.