Funder Experiments,
Experimental Funders
Tom Stafford, Senior Research Fellow
t.stafford@researchonresearch.org
The Research on Research Institute Impact Report (2023-2025)
https://researchonresearch.org/roris-impact-what-weve-achieved-and-where-were-going/
Our definition of experiment
Principled: a research design that allows inference about what causes what (before/after, shadow experiments, true experiment/RCT)
Planned: primary outcome measure and analysis plan declared in advance
Public: a commitment to sharing the results regardless of outcome
These slides:
https://researchonresearch.org/project/a-f-i-r-e/
https://researchonresearch.org/project/a-f-i-r-e/
AFIRE: Accelerator for Funder Experimentation
Sharing work by funders, for funders
Capacity building
Forum
Experiments
Sprints on AI/ML in reviewer selection
Distributed Peer Review, Partial Randomisation, Desk Rejection, and more!
Forum
13 May 16:00 CEST- AI in Funding
17 June 16:00 CEST - Distributed Peer Review
14 October 16:00 CEST - Desk Rejection
16 December 09:00 CEST - ‘Lottery First’ methods
Experimental Funders Group, 2026
Register!
Capacity building
These slides:
The experimental research funder’s handbook (Revised edition, June 2022, ISBN 978-1-7397102-0-0). https://doi.org/10.6084/m9.figshare.19459328.v2
We can plan/run experiments
These slides:
The experimental research funder’s handbook (Revised edition, June 2022, ISBN 978-1-7397102-0-0). https://doi.org/10.6084/m9.figshare.19459328.v2
Experiments
Partial Randomisation
Trials Catalogue
Desk Rejection Shadow Experiment
Can agency staff predict those proposals with the least likelihood of success?
Enquiries:
Josie Coburn,
Research Fellow in Metascience,
Research on Research Institute
josie.coburn@ucl.ac.uk
Evaluating Distributed Peer Review at the Volkswagen Foundation��Anna Butters, Melanie Benson Marshall, Tom Stafford & Stephen Pinfield (Research on Research Institute and University of Sheffield);�Hanna Denecke, Alexander Bondarenko, Barbara Neubauer, Robert Nuske & Pierre Schwidlinski (Volkswagen Foundation)
Distributed Peer Review
Research Questions
1. Can language models help match proposals to reviewers?
2. Is it feasible for something like a conference to adopt/adapt this technology?
3. Can it be done securely/privacy respecting?
Maybe - evidence for meaningful improvements beyond human matching
�Definitely yes
Definitely yes
AI reviewer matching@Metascience2025
Researcher Attitudes to Funder Expts
Where do researchers want innovation?
What should their motives be?
What makes experiments more or less acceptable?
Becoming experimental
Get in touch!��researchonresearch.org�@RoRInstitute�
END
(reserve slides follow)
Other AFIRE experiments
“The Art of the soluble”
Feasible
Interesting
!
Thinking about experiments
Good outcome measures
A part of the fresco “Triumph of Galatea,” created by Raphael around 1512 for the Villa Farnesina in Rome. Art Images via Getty Images
Assays and microscopes
Image: CC Wikimedia
Yes/No
but is it the right question
Close view
but what are you looking for?
Just designing experiments is valuable
The Metascience 2025 conference experiment
Can AI be used for better matching of proposals to reviewers? Feasibility and formal evaluation with the Metascience 2025 conference
Josie Coburn and Tom Stafford 2025-11-26
https://researchonresearch.org/project/a-f-i-r-e/
Finding (enough, good) reviewers is a conceptual and practical problem
Reviewer-proposal matching identified by GRAIL as a key area for possible experiments
Many funders already exploring this
Algorithms need validation!
Meta-metascience
AFIRE Commitment: Observation is not enough - we have to try things!
Metascience 2025 conference, London
The “shadow” experiment
Consent from those submitting and reviewers
All analyses done after final programme decisions
All analyses local - no data left the conference
441 submissions: Title, Abstracts
25 reviewers: assigned to submissions via keywords
1,323 reviews
Research Questions
1. Can language models help match proposals to reviewers?
2. Is it feasible for something like a conference to adopt/adapt this technology?
3. Can it be done securely/privacy respecting?
Matching - via embedding
Reviewer keywords & proposal title+abstract -> embedding space
Code from SNSF: https://github.com/snsf-data/snsf-grant-similarity
Model: SPECTER2: BERT model pre-trained on scientific texts and augmented by a citation graph
You can predict suitability from matching score
…and from this you can predict gain in suitability from using the optimal match
Research Questions
1. Can language models help match proposals to reviewers?
2. Is it feasible for something like a conference to adopt/adapt this technology?
3. Can it be done securely/privacy respecting?
Maybe - evidence for meaningful improvements beyond human matching
�Definitely yes
Definitely yes
Thanks to all participants!