1 of 19

Data Sharing with Endogenous Choices over Differential Privacy Levels

Diptangshu Sen (Georgia Tech)

Joint work with Raef Bassily (OSU), Kate Donahue (MIT/UIUC),

Annuo Zhao, Juba Ziani (Georgia Tech).

16 June 2026

ESIF Economics and AI+ML Meeting || The Econometric Society

Cornell University, Ithaca NY

2 of 19

Motivation: data sharing

………

individually held data

largely insufficient!

share/pool data,

compute together!

Institutions need huge data to build powerful models and computations.

Example: hospitals trying to understand prevalence of new disease in population

Hospital A

Hospital B

Hospital X

sensitive patient data

solution?

3 of 19

Challenges

  1. Privacy protections
    • many tools from differential privacy, but?
    • challengeheterogenous privacy preferences

-> misaligned data sharing incentives among data owners!

  • Push to Decentralize Data Sharing
    • data owners want to retain control over data due to lack of trust
    • challenges – can lead to inefficient sharing!

4 of 19

This talk

  1. Can large data sharing cooperatives be sustained under fully decentralized mechanisms with differential privacy?
  2. Under what conditions?
  3. How efficient are they compared to centralized mechanisms?

5 of 19

Differential Privacy 101

  •  

1-neighboring datasets x, x’

 

 

outcome distributions almost indistinguishable

 

Important property:

Post-processing immunity!

6 of 19

Setting

Target Population

 

 

 

 

 

 

No participation

= no benefits!

 

 

Privacy-accuracy tradeoff!

 

7 of 19

Types of Mechanisms

  •  

 

More Autonomy

More Efficiency

PoS

8 of 19

Accuracy (Variance) of Shared Estimator

  •  

Larger coalitions (higher |S|) better for accuracy!

 

Not clear if large coalitions can be sustained!

9 of 19

Social Cost

  • Each player incurs a burden or cost.
  • Player cost = Accuracy Cost + Privacy Cost

 

 

privacy preference for player ‘i’ (different for different players)

DP attack/observation model: how players perceive risk

 

10 of 19

What is f(S)? (1)

  •  

11 of 19

What is f(S)? (2)

  •  

 

 

 

 

Secure�Aggregation

 

“Get more privacy by hiding in the crowd”

12 of 19

What is f(S)? (3)

  •  

 

Leak (Economic Interpretation): “more people who know about me means more people with strategic power against me”

 

13 of 19

Full Decentralization: Stable Participation

  •  

14 of 19

Flavor of Results (Large ‘n’ regime)

  •  

15 of 19

Comparison

 

 

 

 

 

Centralized

Decentralized

 

 

 

 

 

 

 

 

 

 

PoS

 

 

 

Standard DP

Privacy leakage

Privacy amplfn

16 of 19

Key Takeaways

  • Data sharing incentives depend on the observation model!
  • Full decentralization is bad in general!
  • Useful only when players get strong privacy amplification due to aggregation.
  • Even when it is useful, it is very inefficient in practice!
  • Inefficiency not because of stability requirements, rather

because fully autonomous players choose privacy levels sub-optimally!

17 of 19

Thank You! Questions?

18 of 19

Data Sharing with Endogenous Choices over Differential Privacy Levels

Diptangshu Sen

Joint work with Raef Bassily, Kate Donahue, Annuo Zhao, Juba Ziani.

16 June 2026

ESIF Economics and AI+ML Meeting || The Econometric Society

Cornell University, Ithaca NY

19 of 19

Optimal Stable Coalitions at Large ‘n’

Interested in large coalitions which are:

  • stable;
  • optimal (social cost-minimizing out of all stable ones)

Properties:

  • variance of shared estimator
  • social cost of coalition

When non-trivial compared to no sharing?

 

No Data Sharing