Focus Group Discussion Themes and Reports
LLM4Eval
Getting good at Cranfield
Given a CORPUS and QUERY, a ranker provides a DOCUMENT
With reference to the DOCUMENT and QUERY, a labeller produces a LABEL
We’re getting good at using LLMs here
Where next…?
Different output
Given a CORPUS and QUERY, a ranker provides a DOCUMENT
With reference to the DOCUMENT and QUERY, a labeller produces a LABEL
We’ve traditionally evaluated documents.
What else can we evaluate?
Different “people”
We’ve traditionally represented a searcher just by their query.
We’ve represented relevance by one person’s label.
We’ve not represented interaction well at all.
What can we try instead?
Given a CORPUS and QUERY, a ranker provides a DOCUMENT
With reference to the DOCUMENT and QUERY, a labeller produces a LABEL
Different data
Given a CORPUS and QUERY, a ranker provides a DOCUMENT
With reference to the DOCUMENT and QUERY, a labeller produces a LABEL
We’ve traditionally used crawled documents.
We’ve not done well with private or uncrawlable collections.
Why not try making it up?
Focus Group Discussion Themes and Reports
LLM4Eval
Groups
Group 1: Anti-Personalization
Kenneth Church kenneth.ward.church@gmail.com
Tetsuya Sakai
Key-Sun Choi kschoi@kaist.edu
Mohammad Aliannejadi
Opinions (Anti-Personalization)
Ground News:�Classify news by left v. right
Adding Perspective to Speech/Vision Research�COCO (Objective) 🡪 ArtEmis (Subjective)
BUCC
14
1/20/25
Objective
Facts
Subjective Emotion
More room for Audience Background:
Culture/Language
No Bounding Boxes
Task
Errors of omission and commission
Cultural Sensitivities
Embrace Diversity & Multiple Perspectives
BUCC
19
1/20/25
Hypothesis:
Predictable
Ethics Review
Inter-Annotator Agreement (IAA)
Standard Position
Proposed Alternative
Narrative: DeepSeek Ripped off ChatGPT�(but American bots don’t say the following)
SIGIR-2025
24
7/15/25
Notes
Notes
Group 2: Team anti-anti-personalization
Team Members
Mouly Dewan (mdewan@uw.edu, University of Washington)
Benjamin Vendeville (vendeville@univ-brest.fr, universite bretagne occidentale)
Emma Schüldt (eschuldt@spotify.com, Spotify)
Michael Soprano (michael.soprano@uniud.it, University of Udine)
Jüri Keller (jueri.keller@th-koeln.de, TH Köln)
Poppy Newdick (poppy@spotify.com, Spotify)
Sergio Nunes (sergio.nunes@fe.up.pt, University of Porto)
Jordan Massiah (jormas@amazon.co.uk, Amazon)
Problems that we need evaluation for?
Diversification and Personalization (both ends of the same coin)
Solutions
Efficiency
Problem: How do we efficiently use user signals for evaluation?
Solutions
Alignment
How do we make evaluation metrics that last?
Problems
Solutions
Group 3: BreakEval
��
Group members
Confirmation Bias
Anchoring Bias
Priming Effects
Ordering Effects
Surprises with Human Annotation
Surprises with LLM Judge
Narcissism
Circularity
Benchmark Memorization
Detecting issues with Human Judges
Detecting issues with LLM Judges
Design tests to quantify LLM Evaluation tropes (circularity, Narcissism, Benchmark memorization).
Design content injection attacks.
Human-LLM co-judging
Questions
Creating evaluation decoys for both humans and LLMs to understand their failures. What kind of decoys or methods should we use to make good decoys? How to evaluate these failures? Can we inject scan terms?
Questions
5. How to train LLMs as a judge? Training humans is different from LLMs, with LLMs we try to write better prompts or prompt tuning
6. Are current tasks easy for LLMs? Should we come up with tasks that are complex for LLMs instead of measuring complexity relative to humans?
Conjecture: Maybe do nothing and just wait for a strong LLM which will solve all this!!!
Core Issues
LLM Evaluation Tropes describe 14 things that can go wrong with LLM Judge evaluation
Human Evaluation is taken as a gold standard, but we know there are data quality issues also.
How can we bring Humans + LLMs together, without biasing humans to blindly agree with LLMs?
How can we design tests for “things going wrong”?
Prescription-Strength Judgments
To avoid LLM’s guessing the truth:
How do we get stronger judgments?
Get queries/judges from users that have real information needs. (Similar to Smucker’s study on MovieLens, or OpenAI hiring PhDs to provide human feedback on deep topics)
Dogfooding: design a search engine that IR researchers would like to use for their work.
Work on RAG for complex tasks (long reports, literature surveys)
Group 4:
Bias
People
Qingjing Chen (qingjing.chen@studio.unibo.it)
Kevin Roitero (kevin.roitero@uniud.it)
Gianluca Demartini (demartini@acm.org)
Charles L A Clarke (claclark@gmail.com)
Ronak Pradeep (rpradeep@uwaterloo.ca)
Riccardo Lunardi (riccardo.lunardi@uniud.it)
Maryam Mousavian (m.mousavian@uva.nl)
Stefano Mizzaro (stefano.mizzaro@uniud.it)
44
Minutes (Log)
45
Minutes (Log)
46
Concrete experiment 1 - Gianluca’s data
Possible experiment:
47
Concrete example 2 - Ronak
48
Qingjing Chen
EVENS: Equality versus Equity Notion Spectrum of LLMs, Proceedings of the 20th International Conference on Artificial Intelligence and Law (ICAIL).
49
Group 5:
Evaluation of User Simulation
Group Members
Notes
Evaluation paradigms
Intrinsic: How good the simulator represents the user behavior/action/personality
Extrinsic: Utility for retrieval/information access system
Qualities of simulator to measure
Mark’s mind:
Lets work on evaluating first utterance generation of the model.
e.g., a query generation simulator, but why not just have a human do this for us?
Notes
Notes