1 of 54

Understanding the Bitter Lesson in�Time Series Foundation Models�

Danielle Maddix Robinson

CME 500 SEMINAR

06.02.2026

Senior Applied Scientist, AWS AI

1

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

2 of 54

Collaborators

2

Christos

faloutso@

Annan

annanyu@

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

3 of 54

Time series data

  • Time series are measurements made at regular intervals

3

Energy

Finance

Retail

Weather

Healthcare

Traffic

Time

Value

2024-01-01

2024-01-02

2024-01-03

2024-01-03

2024-01-03

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

4 of 54

Time series forecasting

  • What will happen in the future given the past?

4

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

5 of 54

Probabilistic forecasting

  • Probabilistic forecast captures uncertainty in predictions

5

 

 

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

6 of 54

Local statistical models

  • Fit a separate model for each individual time series
  • Examples: ARIMA, ETS, Theta

  • Strong baseline (esp. limited data)
  • Often interpretable
  • Low flexibility
  • Slow inference

6

 

 

 

 

 

 

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

7 of 54

Global deep learning models

  • Fit a single model for each task
  • Examples: DeepAR, TFT, PatchTST

  • High flexibility
  • Fast inference
  • Slow training
  • Data hungry

7

 

 

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

8 of 54

ChatGPT Moment for Forecasting?

  • Can we develop a single model that both
    • requires no dataset-specific training and
    • performs well on new time series tasks?

8

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

9 of 54

Language modeling and forecasting

9

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

10 of 54

Tokenization

10

Text language models have a discrete vocabulary

Time series are real-valued signals

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

11 of 54

Time series tokenization

11

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

12 of 54

Regression via classification

12

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

13 of 54

Chronos

  • Requires no changes to the language model architecture & training procedure
  • Probabilistic by design

13

✔ probabilistic by design

✔ requires no changes to the language model architecture or training procedure

Ansari, A.F., et al., “Chronos: Learning the Language of Time Series”, TMLR, 2024.

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

14 of 54

Training corpus

14

Language: 15T tokens

Time Series: 84B tokens

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

15 of 54

Training corpus

84B tokens

15

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

16 of 54

Training datasets

  • 28 datasets from various domains and frequencies
  • 890K time series with 84B observations

16

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

17 of 54

Augmenting the training set

TSMix augmentations

Improve pattern diversity by mixing time series from different datasets

Synthetic data

Generate synthetic time series from Gaussian processes

17

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

18 of 54

Baselines

18

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

19 of 54

Evaluation Metrics

19

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

20 of 54

Benchmarks

20

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

21 of 54

Chronos: In-domain Results

21

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

22 of 54

Chronos: Zero-shot Results

22

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

23 of 54

Chronos-Bolt

More accurate and 250x faster than the original Chronos models

23

Chronos-Bolt

Chronos

Input (tokens)

Patches

Individual observations

Output (forecast)

Multi-step quantile forecast

Autoregressive sampling

Loss function

Quantile loss

Cross-entropy loss

Context length

2048

512

Inference device

CPU or GPU

GPU

Ansari, A.F., et al., “Fast and accurate zero-shot forecasting with Chronos-Bolt and AutoGluon”, AWS Technical Report, 2025.

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

24 of 54

Chronos-Bolt⚡: 250x faster than Chronos

24

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

25 of 54

Chronos-Bolt: zero-shot results

25

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

26 of 54

Long-term Behavior on Chaotic Systems

26

Zhang, Y. et al. ”Zero-shot Forecasting of Chaotic Systems," ICLR, 2025.

Zhang, Y. et al., “Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learning”, arXiv preprint arXiv:2505.11349, 2025.

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

27 of 54

Design Choices of TSFMs

27

Yu, A. et al., “Understanding the Implicit Biases of Design Choices for Time Series Foundation Models”, ICLR, 2026.

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

28 of 54

Inductive Biases Overview

28

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

29 of 54

How Do TSFMs Learn Time?

29

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

30 of 54

Temporal Frequency Bias

30

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

31 of 54

Frequency Bias: Good or Bad?

31

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

32 of 54

Temporal Periodicity Bias�

32

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

33 of 54

Periodicity Bias: Good or Bad?

33

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

34 of 54

How Do TSFMs Learn Geometry?

34

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

35 of 54

Design Choice: Embedding Type

Quantization

Continuous Embedding

35

.

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

36 of 54

Geometric Angular Bias�

36

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

37 of 54

Angular Bias: Good or Bad?��

37

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

38 of 54

Geometric Distance Bias

38

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

39 of 54

Distance Bias: Good or Bad?

39

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

40 of 54

Geometric Norm Bias

40

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

41 of 54

Norm Bias: Good or Bad?

41

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

42 of 54

How Do TSFMs Regress to the Mean?

42

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

43 of 54

Regression-to-the-Mean Bias

43

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

44 of 54

Regression-to-the-Mean Bias: Good or bad?

44

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

45 of 54

Understanding Transformers for Time Series

45

Yu, A., Maddix, D.C., et al., ”Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility”, ICLR, 2026.

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

46 of 54

Understanding Transformers for Time Series

46

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

47 of 54

Conclusions

  • Identify design choices in TSFMs that cause 3 inductive biases:
      • Temporal, geometric and regression-to-mean
  • Careful numerical analysis and design of TSFMs is required
  • Temporal data is a different data-modality, e.g., more compressible, has frequency parameters, continuity in time
  • Bitter Lesson
      • Adding traditional forecasting inductive biases can help improve performance on classical benchmarks
      • But it can hurt generalization on unseen domains and tasks,

e.g., chaotic systems

47

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

48 of 54

Chronos-2: From Univariate to Multivariate

48

Ansari, A.F., et al., ”Chronos-2: From Univariate to Universal Forecasting”, arXiv preprint arXiv:2510.15821, 2025.

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

49 of 54

fev-bench: Realistic benchmark for time series forecasting

  • Large-scale evaluation on real-world forecasting tasks
    • 100 univariate & multivariate tasks (incl. 46 with covariates)
  • Statistically sound aggregation methods
    • Reliable model comparisons using bootstrap confidence intervals
  • Extensible infrastructure for reproducible evaluation
    • Lightweight Python wrapper on top of 🤗 datasets library
  • Paper: arxiv.org/abs/2509.26468
  • Code: github.com/autogluon/fev
  • Leaderboard: huggingface.co/spaces/autogluon/fev-bench

49

Shchur, O., et al., ”fev-bench: A Realistic Benchmark for Time Series Forecasting”, arXiv preprint arXiv:2509.26468, 2025.

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

50 of 54

Chronos in the Open Source

  • Inference code available on GitHub

  • Model weights available on Hugging Face 🤗

  • Deploy Chronos-2 on AWS using SageMaker JumpStart

  • Run Chronos with 1 line of code using AutoGluon

50

(Chronos-2 coming soon!)

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

51 of 54

ProbHardE2E

51

Utkarsh, U., Maddix, D.C., Ma, R., Mahoney, M., Wang, Y., "End-to-End Probabilistic Framework for Learning with Hard Constraints”, ICLR 2026.

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

52 of 54

ProbHardE2E

52

Utkarsh, U., Maddix, D.C., Ma, R., Mahoney, M., Wang, Y., "End-to-End Probabilistic Framework for Learning with Hard Constraints”, ICLR 2026.

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

53 of 54

Mitra: Tabular Foundation Model

53

Zhang, X., Maddix, D.C., et al., ”Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models”, NeurIPS, 2025.

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.

54 of 54

Danielle Maddix Robinson

dcmaddix@gmail.com

https://dcmaddix.github.io

54

Thank you!

© 2025, Amazon Web Services, Inc. or its affiliates. All rights reserved.