1 of 30

Funder Experiments,

Experimental Funders

Tom Stafford, Senior Research Fellow

t.stafford@researchonresearch.org

2 of 30

The Research on Research Institute Impact Report (2023-2025)

https://researchonresearch.org/roris-impact-what-weve-achieved-and-where-were-going/

3 of 30

Our definition of experiment

Principled: a research design that allows inference about what causes what (before/after, shadow experiments, true experiment/RCT)

Planned: primary outcome measure and analysis plan declared in advance

Public: a commitment to sharing the results regardless of outcome

4 of 30

These slides:

bit.ly/tomstafford

https://researchonresearch.org/project/a-f-i-r-e/

https://researchonresearch.org/project/a-f-i-r-e/

AFIRE: Accelerator for Funder Experimentation

Sharing work by funders, for funders

Capacity building

Forum

Experiments

Sprints on AI/ML in reviewer selection

Distributed Peer Review, Partial Randomisation, Desk Rejection, and more!

5 of 30

Forum

13 May 16:00 CEST- AI in Funding

17 June 16:00 CEST - Distributed Peer Review

14 October 16:00 CEST - Desk Rejection

16 December 09:00 CEST - ‘Lottery First’ methods

Experimental Funders Group, 2026

Register!

6 of 30

Capacity building

7 of 30

These slides:

bit.ly/tomstafford

The experimental research funder’s handbook (Revised edition, June 2022, ISBN 978-1-7397102-0-0). https://doi.org/10.6084/m9.figshare.19459328.v2

We can plan/run experiments

8 of 30

These slides:

bit.ly/tomstafford

The experimental research funder’s handbook (Revised edition, June 2022, ISBN 978-1-7397102-0-0). https://doi.org/10.6084/m9.figshare.19459328.v2

Experiments

9 of 30

Partial Randomisation

Trials Catalogue

10 of 30

Desk Rejection Shadow Experiment

Can agency staff predict those proposals with the least likelihood of success?

  • a shadow experiment, not an intervention
  • supports optimal use of external review
  • UKRI leading participation
  • recruiting schemes which will complete by end of 2026

Enquiries:

Josie Coburn,

Research Fellow in Metascience,

Research on Research Institute

josie.coburn@ucl.ac.uk

11 of 30

Evaluating Distributed Peer Review at the Volkswagen FoundationAnna Butters, Melanie Benson Marshall, Tom Stafford & Stephen Pinfield (Research on Research Institute and University of Sheffield);�Hanna Denecke, Alexander Bondarenko, Barbara Neubauer, Robert Nuske & Pierre Schwidlinski (Volkswagen Foundation)

Distributed Peer Review

12 of 30

Research Questions

1. Can language models help match proposals to reviewers?

2. Is it feasible for something like a conference to adopt/adapt this technology?

3. Can it be done securely/privacy respecting?

Maybe - evidence for meaningful improvements beyond human matching

�Definitely yes

Definitely yes

AI reviewer matching@Metascience2025

13 of 30

Researcher Attitudes to Funder Expts

Where do researchers want innovation?

What should their motives be?

What makes experiments more or less acceptable?

14 of 30

Becoming experimental

15 of 30

Get in touch!��researchonresearch.org�@RoRInstitute

16 of 30

END

(reserve slides follow)

17 of 30

Other AFIRE experiments

18 of 30

“The Art of the soluble”

Feasible

Interesting

!

Thinking about experiments

19 of 30

Good outcome measures

A part of the fresco “Triumph of Galatea,” created by Raphael around 1512 for the Villa Farnesina in Rome. Art Images via Getty Images

20 of 30

Assays and microscopes

Yes/No

but is it the right question

Close view

but what are you looking for?

21 of 30

Just designing experiments is valuable

22 of 30

The Metascience 2025 conference experiment

23 of 30

Can AI be used for better matching of proposals to reviewers? Feasibility and formal evaluation with the Metascience 2025 conference

Josie Coburn and Tom Stafford 2025-11-26

https://researchonresearch.org/project/a-f-i-r-e/

24 of 30

Finding (enough, good) reviewers is a conceptual and practical problem

  • conceptual: what makes a reviewer good?
  • practical: how do you get a reviewer to agree?

Reviewer-proposal matching identified by GRAIL as a key area for possible experiments

Many funders already exploring this

Algorithms need validation!

25 of 30

Meta-metascience

AFIRE Commitment: Observation is not enough - we have to try things!

  • demonstrate feasibility
  • opportunity for better causal inference

Metascience 2025 conference, London

  • a chance to show we’ll take our own medicine
  • appropriate domain for demonstrating feasibility
  • added value: validate by collecting reviewer self-perception of suitability

26 of 30

The “shadow” experiment

Consent from those submitting and reviewers

All analyses done after final programme decisions

All analyses local - no data left the conference

441 submissions: Title, Abstracts

25 reviewers: assigned to submissions via keywords

1,323 reviews

  • for each we have a match scores & a reviewer suitability judgement
  • (each proposal seen by 3 reviewers)

Research Questions

1. Can language models help match proposals to reviewers?

2. Is it feasible for something like a conference to adopt/adapt this technology?

3. Can it be done securely/privacy respecting?

27 of 30

Matching - via embedding

Reviewer keywords & proposal title+abstract -> embedding space

Code from SNSF: https://github.com/snsf-data/snsf-grant-similarity

  • thanks to Gabriel Okasa and the SNSF data team!

Model: SPECTER2: BERT model pre-trained on scientific texts and augmented by a citation graph

28 of 30

You can predict suitability from matching score

29 of 30

…and from this you can predict gain in suitability from using the optimal match

30 of 30

Research Questions

1. Can language models help match proposals to reviewers?

2. Is it feasible for something like a conference to adopt/adapt this technology?

3. Can it be done securely/privacy respecting?

Maybe - evidence for meaningful improvements beyond human matching

�Definitely yes

Definitely yes

Thanks to all participants!