1
Eduard Gorbunov
June 30, 2025
EUROPT
Samuel Horváth
MBZUAI
Peter Richtárik
KAUST
Martin Takáč
MBZUAI
Nazarii Tupitsa
MBZUAI
Sayantan Choudhury
JHU
Alen Aliev
Skoltech
3
Talk Outline
Motivation
5
Typical Machine Learning Problem
Pictures sources: https://en.wikipedia.org/wiki/Cat
Neural Network
“It is a cat”
“It is a dog”
Goal: classify what is on the picture – cat or dog
(this is my dog)
6
Typical Machine Learning Problem
Empirical Risk Minimization (ERM)
7
Typical Machine Learning Problem
Empirical Risk Minimization (ERM)
8
Typical Machine Learning Problem
Empirical Risk Minimization (ERM)
9
Typical Machine Learning Problem
Empirical Risk Minimization (ERM)
Has a complicated structure
10
Standard Assumptions
Problem:
Convexity:
Smoothness:
11
Standard Assumptions
Problem:
Convexity:
Smoothness:
These assumptions do not hold in Deep Learning!
12
Convexity is not that bad…
Liu et al. Loss landscapes and optimization in over-parameterized non-linear systems and neural networks (Applied and Computational Harmonic Analysis 2022)�Shaipp et al. The surprising agreement between convex optimization theory and learning-rate scheduling for large model training (arXiv:2501.18965)
from (Shaipp et al., 2025)
13
… but smoothness is more problematic
14
… but smoothness is more problematic
15
… but smoothness is more problematic
16
… but smoothness is more problematic
17
Smoothness in Deep Learning
Zhang et al. Why gradient clipping accelerates training: A theoretical justification for adaptivity (ICLR 2020)
Merity et al. Regularizing and optimizing LSTM language models (ICLR 2018)
Zhang et al. (2020) estimated local smoothness constants in the points generated during the training of AWD-LSTM (Merity et al., 2018) on PTB dataset (Clipped-SGD was used)
18
J. Zhang et al. Why gradient clipping accelerates training: A theoretical justification for adaptivity (ICLR 2020)
B. Zhang et al. Improved analysis of clipping algorithms for non-convex optimization (NeurIPS 2020)
Chen et al. . Generalized-smooth nonconvex optimization is as efficient as smooth nonconvex optimization (ICML 2023)
19
J. Zhang et al. Why gradient clipping accelerates training: A theoretical justification for adaptivity (ICLR 2020)
B. Zhang et al. Improved analysis of clipping algorithms for non-convex optimization (NeurIPS 2020)
Chen et al. . Generalized-smooth nonconvex optimization is as efficient as smooth nonconvex optimization (ICML 2023)
20
J. Zhang et al. Why gradient clipping accelerates training: A theoretical justification for adaptivity (ICLR 2020)
B. Zhang et al. Improved analysis of clipping algorithms for non-convex optimization (NeurIPS 2020)
Chen et al. . Generalized-smooth nonconvex optimization is as efficient as smooth nonconvex optimization (ICML 2023)
21
22
23
24
Gradient Descent and Clipping
Gradient Descent (GD)
25
Gradient Descent and Clipping
Gradient Descent (GD)
Clipped Gradient Descent (Clipped-GD)
26
Gradient Descent and Clipping
Gradient Descent (GD)
Clipped Gradient Descent (Clipped-GD)
27
iterations
Li et al. Convex and Non-convex Optimization Under Generalized Smoothness (NeurIPS 2023)
28
Li et al. Convex and Non-convex Optimization Under Generalized Smoothness (NeurIPS 2023)�Koloskova et al. Revisiting Gradient Clipping: Stochastic bias and tight convergence guarantees (ICML 2023)
iterations
iterations
29
iterations
iterations
Li et al. Convex and Non-convex Optimization Under Generalized Smoothness (NeurIPS 2023)�Koloskova et al. Revisiting Gradient Clipping: Stochastic bias and tight convergence guarantees (ICML 2023)
New Results for (Smoothed) Clipping
31
Our ICLR 2025 Paper Resolves This Question
32
Clipped-GD:
33
Clipped-GD:
34
Clipped-GD:
35
Theorem 1
iterations.
36
Theorem 1
iterations.
37
Theorem 1
iterations.
(Li et al., 2023)
(Koloskova et al., 2023)
Li et al. Convex and Non-convex Optimization Under Generalized Smoothness (NeurIPS 2023)�Koloskova et al. Revisiting Gradient Clipping: Stochastic bias and tight convergence guarantees (ICML 2023)
38
39
40
41
42
(“large” gradient regime)
(“small” gradient regime)
43
(“large” gradient regime)
(“small” gradient regime)
44
complexity bound
(“large” gradient regime)
(“small” gradient regime)
Other Extensions
46
Polyak Stepsizes
B. Polyak "Introduction to optimization" (1987)
47
Polyak Stepsizes
Theorem 2
iterations.
B. Polyak "Introduction to optimization" (1987)
48
Polyak Stepsizes
Theorem 2
iterations.
(Takezawa et al., 2024)
B. Polyak "Introduction to optimization" (1987)�Takezawa et al. Parameter-free Clipped Gradient Descent Meets Polyak (NeurIPS 2024)
49
Stochastic Extensions
50
Stochastic Extensions
SGD-PS
51
Stochastic Extensions
SGD-PS
Additional assumptions
52
Stochastic Extensions: Convergence
Theorem 3
53
Adaptive Gradient Descent (AdGD)�(Malitsky & Mishchenko, 2020)
Malitsky & Mishchenko. Adaptive gradient descent without descent (ICML 2020)
54
Adaptive Gradient Descent (AdGD)�(Malitsky & Mishchenko, 2020)
Malitsky & Mishchenko. Adaptive gradient descent without descent (ICML 2020)
55
Adaptive Gradient Descent (AdGD)�(Malitsky & Mishchenko, 2020)
Malitsky & Mishchenko. Adaptive gradient descent without descent (ICML 2020)
56
Adaptive Gradient Descent: New Rate
Theorem 4
57
Adaptive Gradient Descent: New Rate
Theorem 4
58
Adaptive Gradient Descent: New Rate
Theorem 4
59
Acceleration
Gasnikov & Nesterov. Universal fast gradient method for stochastic composite optimization problems (arXiv:1604.05275)
Similar Triangles Method (STM)�(Gasnikov & Nesterov, 2016)
and
60
Acceleration
Gasnikov & Nesterov. Universal fast gradient method for stochastic composite optimization problems (arXiv:1604.05275)
Similar Triangles Method (STM)�(Gasnikov & Nesterov, 2016)
and
61
Acceleration: Convergence Rate
Theorem 5
62
Acceleration: Convergence Rate
Theorem 5
63
Acceleration: Convergence Rate
Theorem 5
64
Numerical Experiments
Logistic regression
Quartic function
65
Numerical Experiments
Logistic regression
Quartic function
Conclusion
67
Main Takeaways
Few Words about MBZUAI
69
MBZUAI Statistics
Computer �Science
Computer �Vision
Machine�Learning
Natural�Language�Processing
Robostics
Statistics and�Data Science
Computational�Biology
Human-Computer�Interaction
70
MBZUAI Statistics
Computer �Science
Computer �Vision
Machine�Learning
Natural�Language�Processing
Robostics
Statistics and�Data Science
Computational�Biology
Human-Computer�Interaction
I will join the StatDS department as an Assistant Professor on August 1st, 2025
71
Open Positions
Samuel Horváth
Assistant Professor (ML)
Martin Takáč�Deputy Department Chair (ML)
Associate Professor
Mladen Kolar
Department Chair (StatDS)�Visiting Professor
Eric Moulines�Professor (ML)
72
ICOMP 2025 Conference & Summer School�(International Conference on Computational Optimization)
Thank you!
eduard.gorbunov@mbzuai.ac.ae
eduardgorbunov.github.io
My email:
Webpage: