Hallucination or Creativity:
How to Evaluate AI-Generated Scientific Stories?
Alex Argese, Pasquale Lisena, Raphaël Troncy
Text2Story @ ECIR 2026
Delft, 29 March 2026
AI Scientist Storyteller
2
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
3
A good scientific story must preserve the paper’s meaning while adapting structure and tone to a target audience.
Factual Summary
The paper introduces the Transformer, a model based on attention mechanisms, removing recurrence and convolution.
Story for General Public
Imagine a reader that doesn’t follow words one by one, but looks at the whole sentence at once, focusing on what matters most. That’s the idea behind the Transformer.
Why evaluating scientific stories is challenging?
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Paper
“Attention Is All You Need” Vaswani et al., 2017
Example
4
Our Contribution
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
5
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Composite evaluation metric
StoryScore
6
BERTScore
(semantic similarity)
Scale: 0–1
High value = story semantically close to the paper
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Composite evaluation metric
StoryScore
7
Prompt Cleanliness
(Prompt Leakage Control)
Scale: 0–1
High value = correct constraint following
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Composite evaluation metric
StoryScore
8
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Composite evaluation metric
StoryScore
Title Coverage
(structural consistency)
Binary scale: 0 or1
Score = 1 if all 5 match exactly → 0 otherwise
9
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Composite evaluation metric
StoryScore
10
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Composite evaluation metric
StoryScore
11
Our Contribution
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Why it matters?
12
Hallucination Detection in Scientific Storytelling
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Why it is hard in a creative setting?
13
Hallucination Detection in Scientific Storytelling
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Story
“It acts like a messenger passing information forward.”
Paper
“The model transfers information between modules.»
Example A - acceptable creative reformulation
Not a hallucination
metaphor, but meaning remains supported
Story
“The method was developed at the University of Birmingham.”
Paper
Anonymized affiliation
Example B - real hallucination
Hallucination
invented factual entity
14
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Hallucination detection�Methods Considered
Evaluation examples
Evaluation examples
15
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Hallucination detection�Methods Considered
Evaluation examples
16
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Hermes, the messenger god of ancient Greece, was known for his speed and efficiency. Similarly, the HERMES system acts as a swift messenger between the initial prompt and the final, refined medical image segmentation.
Evaluation examples
Step 1 — Entity extraction with SpaCy
Step 2 — Semantic support check�For each sentence:
Step 3 — Hallucination rule�A token is hallucinated only if:
The proposed solution, is an automated framework designed to enhance the accuracy of flash memory (FM)-based segmentation.
→ it is a consistent creative simile. Incorrect detection.
→ In the paper, FM stands for Foundation Model, not Flash Memory. Incorrect detection.
Hallucination detection�Methods Considered
17
Need a metric for detecting hallucinations in creative storytelling tasks in the world of scientific research.
↓
Kept the simplest and most effective detector (NER) with SpaCy
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Hallucination detection�Methods Considered
METHOD | Capitalised Words | SpaCy NER | MIRAGE | LLM-Judge (Qwen 7B) | LLM-Judge (GPT 5.1) | HHD (hybrid) |
WHAT IT DETECTS | Surface-form mismatch | Incorrect PERSON/ORG entities | Rewrite-consistency instability | Factual consistency | High-level reasoning errors | Entity + retrieval alignment |
KEY WEAKNESS | Flags abbreviations & creative capitalisations as errors | Misses conceptual errors (wrong claims, invented datasets) | Penalises analogies and audience-adapted rephrasing | "Hallucinates hallucination" labels correct facts as errors | Overcautious: flags benign contextual expansions | Dominant false positives; threshold unstable |
VERDICT | Too noisy | ✓ Chosen | Too rigid | Unstable | Too strict | Unreliable |
18
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Preliminary Evaluation on a story set
19
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
Conclusion and Future Works
20
Thank you for your attention
Text2Story @ ECIR 2026
Delft, 29 March 2026
Alex Argese, Pasquale Lisena, Raphaël Troncy
20
Repository�bit.ly/ai-sci-storyteller �
Mail�alex.argese@eurecom.fr�