1 of 10

TripTide: A Benchmark for Adaptive Travel Planning under Disruptions

Priyanshu Karmakar1, Soumyabrata Chaudhuri1, Shubhojit Mallick2, Manish Gupta2, Abhik Jana1, Shreya Ghosh1

1IIT Bhubaneswar, India 2Microsoft, India

1

{a24cs08008,abhikjana,shreya}@iitbbs.ac.in; chaudhurisoumyabrata@gmail.com

{shubhojit.mallick,gmanish}@microsoft.com

2 of 10

What is TripTide?

Priyanshu Karmakar, Soumyabrata Chaudhuri, Shubhojit Mallick, Manish Gupta, Abhik Jana, Shreya Ghosh. TripTide: A Benchmark for Adaptive Travel Planning under Disruptions. ACL 2026.

  • Benchmark to evaluate how LLMs respond to travel disruptions
  • 3 severity levels: step, day, plan. Traveler tolerance profiles: Flexi-Venturer or Plan-Bound.
  • User Personas: Adventure Seeker, Cultural Explorer, Economical Traveler, or Mountain Enthusiast

2

3 of 10

How are multi-level disruptions handed in TripTide?

Priyanshu Karmakar, Soumyabrata Chaudhuri, Shubhojit Mallick, Manish Gupta, Abhik Jana, Shreya Ghosh. TripTide: A Benchmark for Adaptive Travel Planning under Disruptions. ACL 2026.

3

4 of 10

How is TripTide dataset curated?

Priyanshu Karmakar, Soumyabrata Chaudhuri, Shubhojit Mallick, Manish Gupta, Abhik Jana, Shreya Ghosh. TripTide: A Benchmark for Adaptive Travel Planning under Disruptions. ACL 2026.

  • Disruption Generation: For each of 1000 TripCraft plans, generate 3 disruptions per day using GPT-4o and Gemini 2.5 Pro.
  • Human annotators choose the most contextually meaningful disruption for each plan wrt realism, diversity and traveler-specific relevance.
  • Prompt GPT-4o with orig plan and disruption query to generate a persona-aware revised plan.
  • Automated Script-Based Verification: (1) Valid entities (2) Logical consistency across time, location, and activity sequences.

4

5 of 10

How do we evaluate revised plans?

  • Inputs: current plan, disruption, its severity, and tolerance
  • Output: alternative feasible plan.
  • Preservation of Intent of revised plan
    • Commonsense and hard constraint pass rates (CPR and HCPR)
    • Delivery rate
    • Final pass rate
  • Responsiveness: ratio of fraction of total plans that were mitigated.
  • Adaptability: quantify the semantic, spatial, and sequential shifts between the original and revised itineraries.
    • Semantic closeness between PoIs in the initial and revised itineraries with the user persona using BERT-based cos sim.
    • Diff in spatial convenience (dist to nearest public transit) between the original and revised plans.
    • Changes in order of PoIs across days using normalized edit distance

5

6 of 10

How do LLMs perform on TripTide?

  • GPT-4o is the most stable overall
    • near-perfect delivery rates
    • highest final pass rates
    • low semantic, spatial, and sequential drift
  • Performance degrades mildly with longer horizons

Priyanshu Karmakar, Soumyabrata Chaudhuri, Shubhojit Mallick, Manish Gupta, Abhik Jana, Shreya Ghosh. TripTide: A Benchmark for Adaptive Travel Planning under Disruptions. ACL 2026.

  • Qwen2.5-7B-Instruct
    • Matches GPT-4o on delivery rate and semantic adaptability for shorter plans
    • HCPR and FPR degrade with plan length
  • Phi-4-mini-Instruct
    • Delivery and responsiveness rates improve with longer plans
    • But CPR/HCPR, final pass rates, and adaptability collapse

6

7 of 10

LLM-as-a-Judge based Evaluation

  • Judge revised itinerary from GPT-4o using few-shot Llama-3.1-8B-Instruct
  • 1-5 scale
    • 1 for minimal or ineffective revision
    • 5 for complete resolution of disruption

Priyanshu Karmakar, Soumyabrata Chaudhuri, Shubhojit Mallick, Manish Gupta, Abhik Jana, Shreya Ghosh. TripTide: A Benchmark for Adaptive Travel Planning under Disruptions. ACL 2026.

  • 3-day plans are judged strongest overall
    • highest mean
    • majority Good
    • noticeable number of Excellent samples

7

8 of 10

Human Evaluation of GPT-4o Revised Plans

  • 3-day: edits remained compact with limited downstream consequences
  • 5-day: scope drift (additional edits not strictly required by the disruption)
  • 7-day: widest variance.
    • several good repairs
    • but also spatial or temporal slippage.

Priyanshu Karmakar, Soumyabrata Chaudhuri, Shubhojit Mallick, Manish Gupta, Abhik Jana, Shreya Ghosh. TripTide: A Benchmark for Adaptive Travel Planning under Disruptions. ACL 2026.

8

Semantic issues for a 5-day plan

9 of 10

Human Evaluation of GPT-4o Revised Plans

  • Where the Planner did well
    • Smart swaps
    • Human factors (lighter follow-ups or buffer time)
    • Ground transport Logistics for larger parties
    • Persona fit
  • Where the Planner struggled
    • Superficial fixes: Swapping to a venue with similar limitations.
    • Missing the root cause: Pushed a meal without rebalancing the next day.
    • Real-world timing: Substituted a closed venue with another that was also unavailable at the proposed time.
    • No proper ripple effects

Priyanshu Karmakar, Soumyabrata Chaudhuri, Shubhojit Mallick, Manish Gupta, Abhik Jana, Shreya Ghosh. TripTide: A Benchmark for Adaptive Travel Planning under Disruptions. ACL 2026.

9

Spatial Preservation for a 3 day plan

10 of 10

Summary

Priyanshu Karmakar, Soumyabrata Chaudhuri, Shubhojit Mallick, Manish Gupta, Abhik Jana, Shreya Ghosh. TripTide: A Benchmark for Adaptive Travel Planning under Disruptions. ACL 2026.

  • TripTide: first benchmark for evaluating LLMs’ ability to revise itineraries under realistic disruptions
  • Evaluate
    • Automatic metrics measure Preservation of Intent, Responsiveness, and Adaptability (semantic, spatial, and sequential)
    • LLM-as-a-Judge evaluation
    • Human study assessing revision quality
  • LLMs largely preserve semantic and sequential structure.
  • Disruption handling perf declines as itinerary length increases.

10