1 of 10

HistoryBankQA: Multilingual Temporal Question Answering on Historical Events

Biswadip Mandal1, Anant Khandelwal2, Manish Gupta2

1Amazon US 2Microsoft, India

1

biswadip.iitb@gmail.com, anantk@microsoft.com, gmanish@microsoft.com

2 of 10

Why is temporal reasoning over historical events important?

  • Event extraction: identifying and timestamping events
  • Temporal QA: “Who was the US president during WWII?”
  • Timeline generation: organize events chronologically to build coherent narratives
  • Historical entity linking: disambiguate mentions like “King George” using temporal cues.
  • Event clustering: group events by temporal proximity or shared context
  • Timeline summarization: condense complex narratives, such as social media histories into structured, time-ordered summaries.
  • Temporal NLI: whether one event entails or contradicts another.
  • Benchmarking challenges
    • Lack of multilingual historical event corpus spanning diverse temporal contexts across cultures and periods
    • Lack of a benchmark to evaluate tasks like temporal entailment, ordering, duration inference, and temporal QA

Biswadip Mandal, Anant Khandelwal, Manish Gupta. HistoryBankQA: Multilingual Temporal Question Answering on Historical Events. STARSEM 2026.

2

3 of 10

HistoryBank: A Large-scale Multilingual DB of Historical Events

  • Multilingual DB of 10M+ events sourced from Wikipedia OnThisDay and infoboxes.
  • 10 langs: en, bn, de, fr, id, hi, it, pt, ru, es.
  • Domains: politics, culture, science, sports.
  • Temporal KV pairs from Infobox 🡪 NL events using GPT4o.
  • Wikipedia’s On This Day pages: Births, deaths and generic events
  • Annotate relevant countries, and country-specific polarity.
  • Prepend article title to each event description.
  • 12K events dated before 0 (BC).

Biswadip Mandal, Anant Khandelwal, Manish Gupta. HistoryBankQA: Multilingual Temporal Question Answering on Historical Events. STARSEM 2026.

3

4 of 10

HistoryBankQA: Temporal QA Benchmark

  • Joint evaluation of memorization (factual recall) and reasoning (ordering, comparison, disambiguation)
  • Multilingual: Temporal expressions, date formats, aspect, tense, and event-ordering cues differ across langs.
  • FactQA: Generate question using an LM.
  • DurationQA: Sort events for an article, form start-end event pairs, filter valid pairs with a LM classifier, ask LM to create questions

Biswadip Mandal, Anant Khandelwal, Manish Gupta. HistoryBankQA: Multilingual Temporal Question Answering on Historical Events. STARSEM 2026.

  • RecurrenceQA: Get me year of previous occurrence of recurrent events (elections, sports leagues, tournaments)
  • RelationQA: Over pairs of event spans
  • SequenceQA: Sequence Arrangement, MCQ, Order Verification (True/False)
  • CountQA: Count Between Events, Count Between Years, Count by Century

4

5 of 10

Examples from HistoryBankQA

Biswadip Mandal, Anant Khandelwal, Manish Gupta. HistoryBankQA: Multilingual Temporal Question Answering on Historical Events. STARSEM 2026.

5

6 of 10

Examples from HistoryBankQA

Biswadip Mandal, Anant Khandelwal, Manish Gupta. HistoryBankQA: Multilingual Temporal Question Answering on Historical Events. STARSEM 2026.

6

7 of 10

How does GPT4o perform on HistoryBankQA?

  • 10,000 test examples from each task per language.
  • Test data with 536K samples across all languages.
  • Zeroshot GPT4o exact match accuracy
  • Best for en.
  • Between_events shows lowest perf for CountQA
  • SequenceQA: MCQ, verify, arrange (easy to hard)

Biswadip Mandal, Anant Khandelwal, Manish Gupta. HistoryBankQA: Multilingual Temporal Question Answering on Historical Events. STARSEM 2026.

7

8 of 10

How do LLMs perform on HistoryBankQA?

  • LLaMA-3-8B, Mistral7B, Gemma-2-9B, Qwen3-8B, and GPT4o.
  • Despite being relatively low-resource compared to en/de, hi and id have high perf.
  • GPT-4o maintains relatively stable multilingual perf; smaller models show pronounced language-specific degradation.
  • Models perform well for RecurrenceQA and DurationQA.

Biswadip Mandal, Anant Khandelwal, Manish Gupta. HistoryBankQA: Multilingual Temporal Question Answering on Historical Events. STARSEM 2026.

  • GPT-4o is best; Gemma2 leads among smaller models.
    • Gemma > Qwen > Mistral > LLaMA
  • RelationQA
    • Lowest accuracy: Gap Between, Overlap, and Duration Difference
    • Demand a multi-step reasoning process
    • Highest accuracy: Order Span Start

8

9 of 10

Human Evaluation

  • 100 randomly sampled English events
  • Faithfulness: generated event desc reflects infobox info?
    • 97% yes; 3% minor issues.
  • Event Interpretability: is event understandable as a standalone historical event?
    • 88% yes, 6% partially clear motivating us to add Wiki title to event desc (e.g., “The television show last aired”)
  • Event Interestingness: discovered event corresponds to a real-world event?
    • 84% yes; 16% are administrative mentions, routine statistical updates, or census records (e.g., population counts).

Biswadip Mandal, Anant Khandelwal, Manish Gupta. HistoryBankQA: Multilingual Temporal Question Answering on Historical Events. STARSEM 2026.

  • Error Analysis for en GPT4o
    • Tasks needing multiple reasoning steps
    • Lack or misrepresent the required temporal knowledge, resulting in incorrect event dates
    • Flawed reasoning leading to wrong temporal comparisons, inconsistent boundary handling or misordered events.
    • Reasoning-answer mismatches lead to wrong final answer contradicting the stated reasoning.

9

10 of 10

Summary

  • HistoryBank: event DB with 10M+ events across 10 languages, built from Wikipedia’s On This Day pages and event-centric infoboxes.
  • HistoryBankQA: 6 multilingual temporal QA tasks to test factual recall and reasoning.
  • Zero shot GPT-4o and Gemma2 are best. Lot of room for improvement.
  • RAG did not work.

10