1 of 20

Hallucination or Creativity:

How to Evaluate AI-Generated Scientific Stories?

Alex Argese, Pasquale Lisena, Raphaël Troncy

Text2Story @ ECIR 2026

Delft, 29 March 2026

2 of 20

AI Scientist Storyteller

  • Turn scientific papers into stories
  • Adapt to different personas:
    • Researchers & Engineers
    • Student
    • Teacher
    • Policy Maker
    • Jounalist
    • Investor
    • General pulic

2

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

3 of 20

3

A good scientific story must preserve the paper’s meaning while adapting structure and tone to a target audience.

Factual Summary

The paper introduces the Transformer, a model based on attention mechanisms, removing recurrence and convolution.

Story for General Public

Imagine a reader that doesn’t follow words one by one, but looks at the whole sentence at once, focusing on what matters most. That’s the idea behind the Transformer.

Why evaluating scientific stories is challenging?

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Paper

“Attention Is All You Need” Vaswani et al., 2017

Example

  • Traditional metrics alone (ROUGE, BERTScore, etc.) are not enough
  • We want stories that are factual accurate, but have freedom in structure, for engagement and creativity

4 of 20

  • A new metric for AI-generated scientific stories:�the StoryScore metric

  • A comparison of different hallucination detectors for AI-generated scientific stories

4

Our Contribution

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

5 of 20

5

 

 

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Composite evaluation metric

StoryScore

6 of 20

6

 

BERTScore

(semantic similarity)

  • Measures semantic alignment with the source paper

  • BERTScore between concatenated story sections and paper context

  • Model: roberta-large

Scale: 0–1

High value = story semantically close to the paper

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Composite evaluation metric

StoryScore

7 of 20

7

 

Prompt Cleanliness

(Prompt Leakage Control)

  • Verifies absence of prompt or control artefacts in the output

  • Flags residual instructions, schema markers, or malformed structure

Scale: 0–1

High value = correct constraint following

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Composite evaluation metric

StoryScore

8 of 20

8

 

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Composite evaluation metric

StoryScore

Title Coverage

(structural consistency)

  • Comparison after normalisation�of the Storyteller titles and Splitter titles.

  • Necessary to check the consistency of the structure

Binary scale: 0 or1

Score = 1 if all 5 match exactly → 0 otherwise

9 of 20

9

 

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Composite evaluation metric

StoryScore

 

10 of 20

10

 

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Composite evaluation metric

StoryScore

 

11 of 20

  • A new metric for AI-generated scientific stories:�the StoryScore metric

  • A comparison of different hallucination detectors for AI-generated scientific stories

11

Our Contribution

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

12 of 20

Why it matters?

  • A scientific story must remain faithful to the paper
  • Invented entities, wrong names, or unsupported claims reduce trust
  • Hallucination detection is therefore a key evaluation dimension

12

Hallucination Detection in Scientific Storytelling

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Why it is hard in a creative setting?

  • Stories may use metaphors and analogies
  • Stories may simplify terms for non-expert audiences
  • Not every wording difference is a hallucination

13 of 20

13

Hallucination Detection in Scientific Storytelling

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Story

“It acts like a messenger passing information forward.”

Paper

“The model transfers information between modules.»

Example A - acceptable creative reformulation

Not a hallucination

metaphor, but meaning remains supported

Story

“The method was developed at the University of Birmingham.”

Paper

Anonymized affiliation

Example B - real hallucination

Hallucination

invented factual entity

14 of 20

14

 

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Hallucination detection�Methods Considered

  • MIRAGE*
    • MIRAGE re-generates multiple rewrites of the story
    • Computes alignment between the original generation and its own rewrites
    • If a concept is unstable or unsupported, the model assigns high hallucination probability
  • “University of Birmingham” → not in paper correctly detected

  • story uses “AI”, paper uses “Artificial Intelligence” → incorrectly flagged

Evaluation examples

  • to explain what is Robustness in AI, the story used a simpler example about object detection in images�MIRAGE considers the metaphor as a hallucination

Evaluation examples

15 of 20

15

  • LLM-as-a-judge (Qwen 7B & GPT 5.1)
    • A Large model (Qwen-7B) receives:
      • CONTEXT: the paper text
      • ANSWER: the story section

    • Generates a JSON containing:
      • name_accuracy
      • numeric_accuracy
      • overall_faithfulness
      • hallucinated_names / numbers

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Hallucination detection�Methods Considered

  • "Robust-kit, developed at the University of Birmingham…»�hallucination NOT detected

  • Correct story (human checked) LLM-as-judge incorrectly outputs hallucinated names:�["SwinUnetR", "MedSAM", "SAM-Med2D"]�these appear in the paper but not in that section�→ incorrect flagged

Evaluation examples

16 of 20

16

  • Hybrid Hallucination Detection (HHD)

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Hermes, the messenger god of ancient Greece, was known for his speed and efficiency. Similarly, the HERMES system acts as a swift messenger between the initial prompt and the final, refined medical image segmentation.

Evaluation examples

Step 1 — Entity extraction with SpaCy

  • Identify “technical tokens”: Capitalised words, acronyms, numbers…
  • Ignore metaphorical or generic words → no false positives

Step 2 — Semantic support check�For each sentence:

  • Find top-k most similar paper sentences (MiniLM embeddings)
  • Check if each technical token appears in those contexts

Step 3 — Hallucination rule�A token is hallucinated only if:

  • It does NOT appear in any top-k similar paper sentences
  • AND the similarity < 0.6

The proposed solution, is an automated framework designed to enhance the accuracy of flash memory (FM)-based segmentation.

→ it is a consistent creative simile. Incorrect detection.

→ In the paper, FM stands for Foundation Model, not Flash Memory. Incorrect detection.

Hallucination detection�Methods Considered

17 of 20

17

Need a metric for detecting hallucinations in creative storytelling tasks in the world of scientific research.

Kept the simplest and most effective detector (NER) with SpaCy

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Hallucination detection�Methods Considered

METHOD

Capitalised Words

SpaCy NER

MIRAGE

LLM-Judge (Qwen 7B)

LLM-Judge (GPT 5.1)

HHD (hybrid)

WHAT IT DETECTS

Surface-form mismatch

Incorrect PERSON/ORG entities

Rewrite-consistency instability

Factual consistency

High-level reasoning errors

Entity + retrieval alignment

KEY WEAKNESS

Flags abbreviations & creative capitalisations as errors

Misses conceptual errors (wrong claims, invented datasets)

Penalises analogies and audience-adapted rephrasing

"Hallucinates hallucination" labels correct facts as errors

Overcautious: flags benign contextual expansions

Dominant false positives; threshold unstable

VERDICT

Too noisy

✓ Chosen

Too rigid

Unstable

Too strict

Unreliable

18 of 20

18

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Preliminary Evaluation on a story set

  • 76 generated stories (Pre-trained vs Fine-tuned pipeline)
  • Fine-tuning strongly improves overall StoryScore
  • The biggest gain is Prompt Cleanliness: prompt leakage is essentially removed
  • Better human readability confirmed by StoryScore: 0.560 → 0.787

19 of 20

19

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

Conclusion and Future Works

  • These metrics provide approximations rather than ground-truth guarantees

  • Improve grounding mechanisms and hallucination mitigation

  • Scale up evaluation with larger and more diverse evaluator

20 of 20

20

Thank you for your attention

Text2Story @ ECIR 2026

Delft, 29 March 2026

Alex Argese, Pasquale Lisena, Raphaël Troncy

20

Repository�bit.ly/ai-sci-storyteller