1 of 73

1

 

Eduard Gorbunov

June 30, 2025

EUROPT

2 of 73

Samuel Horváth

MBZUAI

Peter Richtárik

KAUST

Martin Takáč

MBZUAI

Nazarii Tupitsa

MBZUAI

Sayantan Choudhury

JHU

Alen Aliev

Skoltech

 

3 of 73

3

Talk Outline

  1. Motivation��
  2. New Results for (Smoothed) Clipping

  • Other Extensions

4 of 73

Motivation

5 of 73

5

Typical Machine Learning Problem

Neural Network

“It is a cat”

“It is a dog”

 

Goal: classify what is on the picture – cat or dog

(this is my dog)

6 of 73

6

Typical Machine Learning Problem

 

Empirical Risk Minimization (ERM)

7 of 73

7

Typical Machine Learning Problem

 

Empirical Risk Minimization (ERM)

8 of 73

8

Typical Machine Learning Problem

 

Empirical Risk Minimization (ERM)

 

 

 

9 of 73

9

Typical Machine Learning Problem

 

Empirical Risk Minimization (ERM)

 

 

 

Has a complicated structure

10 of 73

10

Standard Assumptions

Problem:

Convexity:

Smoothness:

11 of 73

11

Standard Assumptions

Problem:

Convexity:

Smoothness:

These assumptions do not hold in Deep Learning!

12 of 73

12

Convexity is not that bad…

  • Tricks from convex optimization (e.g., momentum) work in DL�
  • Some relaxations of convexity hold for overparameterized models (Liu et al., 2022)�
  • There is a surprising agreement between DL and convex optimization theory (Shaipp et al., 2025)

Liu et al. Loss landscapes and optimization in over-parameterized non-linear systems and neural networks (Applied and Computational Harmonic Analysis 2022)�Shaipp et al. The surprising agreement between convex optimization theory and learning-rate scheduling for large model training (arXiv:2501.18965)

from (Shaipp et al., 2025)

13 of 73

13

… but smoothness is more problematic

  • Not satisfied even for simple neural networks (e.g., linear ones)

14 of 73

14

… but smoothness is more problematic

  • Not satisfied even for simple neural networks (e.g., linear ones)�
  • Does not explain the behavior of the methods in DL

15 of 73

15

… but smoothness is more problematic

  • Not satisfied even for simple neural networks (e.g., linear ones)�
  • Does not explain the behavior of the methods in DL�
  • Even if satisfied, the smoothness constant can depend on the points

16 of 73

16

… but smoothness is more problematic

  • Not satisfied even for simple neural networks (e.g., linear ones)�
  • Does not explain the behavior of the methods in DL�
  • Even if satisfied, the smoothness constant can depend on the points
  • Any alternatives?

17 of 73

17

Smoothness in Deep Learning

Zhang et al. Why gradient clipping accelerates training: A theoretical justification for adaptivity (ICLR 2020)

Merity et al. Regularizing and optimizing LSTM language models (ICLR 2018)

Zhang et al. (2020) estimated local smoothness constants in the points generated during the training of AWD-LSTM (Merity et al., 2018) on PTB dataset (Clipped-SGD was used)

18 of 73

18

 

 

J. Zhang et al. Why gradient clipping accelerates training: A theoretical justification for adaptivity (ICLR 2020)

B. Zhang et al. Improved analysis of clipping algorithms for non-convex optimization (NeurIPS 2020)

Chen et al. . Generalized-smooth nonconvex optimization is as efficient as smooth nonconvex optimization (ICML 2023)

19 of 73

19

 

 

  • B. Zhang et al. (2020) generalized the above to just differentiable functions

J. Zhang et al. Why gradient clipping accelerates training: A theoretical justification for adaptivity (ICLR 2020)

B. Zhang et al. Improved analysis of clipping algorithms for non-convex optimization (NeurIPS 2020)

Chen et al. . Generalized-smooth nonconvex optimization is as efficient as smooth nonconvex optimization (ICML 2023)

20 of 73

20

 

 

  • B. Zhang et al. (2020) generalized the above to just differentiable functions
  • Chen et al. (2023) proved that the above is equivalent to

J. Zhang et al. Why gradient clipping accelerates training: A theoretical justification for adaptivity (ICLR 2020)

B. Zhang et al. Improved analysis of clipping algorithms for non-convex optimization (NeurIPS 2020)

Chen et al. . Generalized-smooth nonconvex optimization is as efficient as smooth nonconvex optimization (ICML 2023)

21 of 73

21

 

 

22 of 73

22

 

 

23 of 73

23

 

 

24 of 73

24

Gradient Descent and Clipping

Gradient Descent (GD)

25 of 73

25

Gradient Descent and Clipping

Gradient Descent (GD)

Clipped Gradient Descent (Clipped-GD)

26 of 73

26

Gradient Descent and Clipping

 

Gradient Descent (GD)

Clipped Gradient Descent (Clipped-GD)

27 of 73

27

 

 

iterations

 

Li et al. Convex and Non-convex Optimization Under Generalized Smoothness (NeurIPS 2023)

28 of 73

28

 

 

Li et al. Convex and Non-convex Optimization Under Generalized Smoothness (NeurIPS 2023)�Koloskova et al. Revisiting Gradient Clipping: Stochastic bias and tight convergence guarantees (ICML 2023)

iterations

 

iterations

 

 

29 of 73

29

 

 

iterations

 

iterations

 

 

  • Can we prove better rates without additional assumptions?

Li et al. Convex and Non-convex Optimization Under Generalized Smoothness (NeurIPS 2023)�Koloskova et al. Revisiting Gradient Clipping: Stochastic bias and tight convergence guarantees (ICML 2023)

30 of 73

New Results for (Smoothed) Clipping

31 of 73

31

Our ICLR 2025 Paper Resolves This Question

32 of 73

32

 

Clipped-GD:

33 of 73

33

 

Clipped-GD:

 

34 of 73

34

 

Clipped-GD:

 

 

35 of 73

35

 

Theorem 1

 

iterations.

36 of 73

36

 

Theorem 1

 

iterations.

 

37 of 73

37

 

Theorem 1

 

iterations.

 

(Li et al., 2023)

(Koloskova et al., 2023)

Li et al. Convex and Non-convex Optimization Under Generalized Smoothness (NeurIPS 2023)�Koloskova et al. Revisiting Gradient Clipping: Stochastic bias and tight convergence guarantees (ICML 2023)

38 of 73

38

 

  • Standard start of the proof:

39 of 73

39

 

  • Standard start of the proof:

 

40 of 73

40

 

  • Standard start of the proof:

 

41 of 73

41

 

  • Then, we consider possible cases and prove:

42 of 73

42

 

  • Then, we consider possible cases and prove:

(“large” gradient regime)

(“small” gradient regime)

43 of 73

43

 

  • Then, we consider possible cases and prove:

(“large” gradient regime)

(“small” gradient regime)

 

44 of 73

44

 

  • Then, we consider possible cases and prove:

 

complexity bound

(“large” gradient regime)

(“small” gradient regime)

45 of 73

Other Extensions

46 of 73

46

Polyak Stepsizes

B. Polyak "Introduction to optimization" (1987)

47 of 73

47

Polyak Stepsizes

Theorem 2

 

iterations.

B. Polyak "Introduction to optimization" (1987)

48 of 73

48

Polyak Stepsizes

Theorem 2

 

iterations.

  • Improved complexity compared to

(Takezawa et al., 2024)

B. Polyak "Introduction to optimization" (1987)�Takezawa et al. Parameter-free Clipped Gradient Descent Meets Polyak (NeurIPS 2024)

49 of 73

49

Stochastic Extensions

50 of 73

50

Stochastic Extensions

 

SGD-PS

 

51 of 73

51

Stochastic Extensions

 

SGD-PS

 

Additional assumptions

 

52 of 73

52

Stochastic Extensions: Convergence

Theorem 3

 

53 of 73

53

Adaptive Gradient Descent (AdGD)�(Malitsky & Mishchenko, 2020)

Malitsky & Mishchenko. Adaptive gradient descent without descent (ICML 2020)

54 of 73

54

Adaptive Gradient Descent (AdGD)�(Malitsky & Mishchenko, 2020)

Malitsky & Mishchenko. Adaptive gradient descent without descent (ICML 2020)

 

 

55 of 73

55

Adaptive Gradient Descent (AdGD)�(Malitsky & Mishchenko, 2020)

Malitsky & Mishchenko. Adaptive gradient descent without descent (ICML 2020)

 

 

56 of 73

56

Adaptive Gradient Descent: New Rate

Theorem 4

 

 

57 of 73

57

Adaptive Gradient Descent: New Rate

Theorem 4

 

 

  • Improved rate compared to

58 of 73

58

Adaptive Gradient Descent: New Rate

Theorem 4

 

 

 

59 of 73

59

Acceleration

Gasnikov & Nesterov. Universal fast gradient method for stochastic composite optimization problems (arXiv:1604.05275)

Similar Triangles Method (STM)�(Gasnikov & Nesterov, 2016)

and

60 of 73

60

Acceleration

Gasnikov & Nesterov. Universal fast gradient method for stochastic composite optimization problems (arXiv:1604.05275)

Similar Triangles Method (STM)�(Gasnikov & Nesterov, 2016)

 

and

61 of 73

61

Acceleration: Convergence Rate

Theorem 5

 

62 of 73

62

Acceleration: Convergence Rate

Theorem 5

 

 

63 of 73

63

Acceleration: Convergence Rate

Theorem 5

 

 

64 of 73

64

Numerical Experiments

Logistic regression

Quartic function

 

65 of 73

65

Numerical Experiments

Logistic regression

Quartic function

 

66 of 73

Conclusion

67 of 73

67

Main Takeaways

 

  • Lower bounds?
  • Tightness of the derived results?
  • Adaptive stochastic methods?
  • More general assumptions?
  • Extensions to other problem classes?

68 of 73

Few Words about MBZUAI

69 of 73

69

MBZUAI Statistics

  • The world’s first graduate-level, research-based university solely dedicated to AI�
  • Eight research departments

Computer �Science

Computer �Vision

Machine�Learning

Natural�Language�Processing

Robostics

Statistics and�Data Science

Computational�Biology

Human-Computer�Interaction

70 of 73

70

MBZUAI Statistics

  • The world’s first graduate-level, research-based university solely dedicated to AI�
  • Eight research departments

Computer �Science

Computer �Vision

Machine�Learning

Natural�Language�Processing

Robostics

Statistics and�Data Science

Computational�Biology

Human-Computer�Interaction

I will join the StatDS department as an Assistant Professor on August 1st, 2025

71 of 73

71

Open Positions

  • We have several open positions (jointly with Eric, Martin, Mladen, Samuel) for
  • Interns
  • Postdocs
  • Prospective PhD students – please, apply through MBZUAI website

Samuel Horváth

Assistant Professor (ML)

Martin Takáč�Deputy Department Chair (ML)

Associate Professor

Mladen Kolar

Department Chair (StatDS)�Visiting Professor

Eric Moulines�Professor (ML)

72 of 73

72

ICOMP 2025 Conference & Summer School�(International Conference on Computational Optimization)

  • Where: Abu Dhabi, UAE�
  • When:
  • Summer School: October 14-16, 2025
  • Conference: October 17-19, 2025
  • Deadlines:
  • Abstract: August 1, 2025
  • Full paper: August 15, 2025
  • Keynote speakers: Volkan Cevher, Eric Moulines, Vladimir Spokoiny, Daniel Kunh, Katya Scheinberg, and many others!

73 of 73

Thank you!

eduard.gorbunov@mbzuai.ac.ae

eduardgorbunov.github.io

My email:

Webpage: