Scientific Machine Learning for Modeling, Optimization, and Control with Safety Guarantees
Ján Drgoňa
Associate Professor �Civil and Systems Engineering Department �Electrical and Computer Engineering Department (secondary)�The Ralph O'Connor Sustainable Energy Institute (ROSEI)
Data Science and AI Institute (DSAI)�
2
Why Safety Matters
Nominal conditions → system stable and constraints satisfied.
Why Safety Matters
Nominal conditions → system stable and constraints satisfied.
Plant-model mismatch + disturbance → constraints violation.
Few moments later
Optimizing Complex Systems with Safety Guarantees is Hard
4
Challenge 1: Heterogenous Modeling Methods
5
Data-driven
Physics-based
More domain knowledge
Less domain knowledge
White-box models
Gray-box models
Black-box models
Challenge 2: Heterogenous Solution Methods
Constrained Optimization
Differential Equations
Supervised Learning
Reinforcement Learning
Less domain knowledge
More domain knowledge
Challenge 3: Heterogenous Solution Tools
Constrained Optimization
Differential Equations
Supervised Learning
Reinforcement Learning
More domain knowledge
Less domain knowledge
Automatic Differentiation (AD) in Machine Learning
8
Baydin, Atilim Gunes et al. Automatic differentiation in machine learning: a survey. Journal of Machine Learning Research, 2015
Animation source: wikipedia
Gradient Descent Algorithm
Backpropagation Algorithm
AD enables efficient and accurate gradient computation, which is fundamental for training complex ML models using GPUs.
SW and HW Innovations
What?
9
Why?
How?
Image source: https://sciml.wur.nl/reviews/sciml/sciml.html
Karniadakis, G.E., Kevrekidis, I.G., Lu, L. et al. Physics-informed machine learning. Nat Rev Phys 3, 422–440, 2021.
Scientific Machine Learning (SciML)
Selected Scientific Machine Learning Literature
10
Components of Scientific Machine Learning
11
Karniadakis, G.E., Kevrekidis, I.G., Lu, L. et al. Physics-informed machine learning. Nat Rev Phys 3, 2021.
Thiyagalingam, J., Shankar, M., Fox, G. et al. Scientific machine learning benchmarks. Nature Reviews Physics 4, 413–420, 2022.
Nghiem T., Drgona J., et al. Physics-Informed Machine Learning for Modeling and Control of Dynamical Systems, ACC, 2023.
Learning to Solve Differential Equations with Physics-Informed Neural Networks (PINNs)
12
Training neural networks as PDE solutions
Application: Parameter estimation from data
M. Raissi, et al., Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations, Journal of Computational Physics, 2019
Images: NVIDIA Modulus
Learning to Solve Differential Equations with Physics-Informed Neural Networks (PINNs)
13
M. Raissi, et al., Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations, Journal of Computational Physics, 2019
Dataset: collocation points in the spatio-temporal coordinates.
Architecture: PDE equations solved with neural network via automatic differentiation.
Loss function: minimizing PDE equation, initial and boundary condition residuals.
Learning to Optimize (L2O) with Constraints
14
Training neural networks as optimization solutions
Application: solving optimal power flow
James Kotary, et al., End-to-End Constrained Optimization Learning: A Survey, IJCAI, 2021
Learning to Optimize (L2O) with Constraints
15
A. Agrawal, et al., Differentiable Convex Optimization Layers, 2019
P. Donti, et al., DC3: A learning method for optimization with hard constraints, 2021
Dataset: collocation points in the parametric space.
Loss function: minimizing objective function and constraints penalties.
Architecture: differentiable optimization solver with neural network surrogate.
16
Nonlinear system identification
R. T. Q. Chen, et al., Neural ordinary differential equations. NeurIPS, 2018
C. Rackauckas, et al., Universal Differential Equations for Scientific Machine Learning, 2021
James Koch, et al., Learning Neural Differential Algebraic Equations via Operator Splitting, CDC, 2025
Applications: modeling process dynamics
Learning to Model (L2M) Dynamical Systems
17
R. T. Q. Chen, et al., Neural Ordinary Differential Equations, 2019
B. Lusch, et al., Deep learning for universal linear embeddings of nonlinear dynamics, 2018
Dataset: time-series of states, inputs, and disturbances tuples.
Loss function: trajectory matching, regularizations, and constraints penalties.
Architecture: differentiable ODE solver with neural network model.
Architecture: Koopman operator with neural network basis functions.
Learning to Model (L2M) Dynamical Systems
Learning to Control (L2C) Methodologies
18
Supervised L2C: Approximate Model Predictive Control
Self-Supervised L2C: Differentiable Predictive Control (DPC)
J. Drgoňa, A. Tuor and D. Vrabie, "Learning Constrained Parametric Differentiable Predictive Control Policies With Guarantees," in IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2024
Ján Drgoňa, et al, Differentiable predictive control: Deep learning alternative to explicit model predictive control for unknown nonlinear systems, Journal of Process Control, 2022
M. Hertneck, et al., "Learning an Approximate Model Predictive Controller With Guarantees," in IEEE Control Systems Letters, 2018
B. Karg and S. Lucia, "Efficient Representation and Approximation of Model Predictive Control Laws via Deep Learning," in IEEE Transactions on Cybernetics, 2020
Supervised L2C: Approximate MPC
19
Step 1: solve set of MPC problems to generate labeled training data
Step 2: supervised imitation learning to learn approximate MPC policy
M. Hertneck, et al., "Learning an Approximate Model Predictive Controller With Guarantees," in IEEE Control Systems Letters, 2018
Supervised L2C: Approximate MPC
20
Step 1: solve set of MPC problems to generate labeled training data
Step 2: supervised imitation learning to learn approximate MPC policy
M. Hertneck, et al., "Learning an Approximate Model Predictive Controller With Guarantees," in IEEE Control Systems Letters, 2018
Problem 1: data generation is expensive!
Problem 2: hard to integrate constraints into supervised learning.
Self-Supervised L2C: Differentiable Predictive Control (DPC)
21
J. Drgoňa, A. Tuor and D. Vrabie, "Learning Constrained Parametric Differentiable Predictive Control Policies With Guarantees," in IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2024
Dataset: collocation points in the control parametric space.
Loss function: reference tracking, constraints and terminal penalties.
Architecture: differentiable model with neural network control policy.
Differentiable Closed-Loop System
22
Differentiable Closed-Loop System
23
DPC Policy Optimization Algorithm
24
DPC vs Model-based Reinforcement Learning
25
J. Drgoňa, A. Tuor and D. Vrabie, Learning Constrained Parametric Differentiable Predictive Control Policies With Guarantees, in IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2024
J. Drgoňa, et al., Differentiable Predictive Control: An MPC Alternative for Unknown Nonlinear Systems using Constrained Deep Learning, Journal of Process Control, 2022
DPC is closely related to MBRL in that both leverage model of dynamics, but DPC utilizes differentiable closed-loop models and cost functions, allowing direct policy gradients without requiring a learned critic.
Empirical risk minimization problem:
Policy gradient via automatic differentiation:
Critic parametrized by differentiable MPC loss function:
DPC vs Model Predictive Control
26
J. Drgoňa, A. Tuor and D. Vrabie, Learning Constrained Parametric Differentiable Predictive Control Policies With Guarantees, in IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2024
J. Drgoňa, et al., Differentiable Predictive Control: An MPC Alternative for Unknown Nonlinear Systems using Constrained Deep Learning, Journal of Process Control, 2022
DPC is also related to explicit MPC, as it solves a parametric optimal control problem via gradient-based policy optimization.
There is a structural equivalence between single shooting formulation of MPC problem and the unrolled closed-loop system dynamics in DPC.
DPC is also closely related to MPC, but instead of single instance online optimization, the DPC learns parametric explicit policy offline, over distribution of parametric instances in batched setting.
DPC vs Model Predictive Control
27
J. Drgoňa, A. Tuor and D. Vrabie, Learning Constrained Parametric Differentiable Predictive Control Policies With Guarantees, in IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2024
J. Drgoňa, et al., Differentiable Predictive Control: An MPC Alternative for Unknown Nonlinear Systems using Constrained Deep Learning, Journal of Process Control, 2022
DPC is also related to explicit MPC, as it solves a parametric optimal control problem via gradient-based policy optimization.
There is a structural equivalence between single shooting formulation of MPC problem and the unrolled closed-loop system dynamics in DPC.
DPC is also closely related to MPC, but instead of single instance online optimization, the DPC learns parametric explicit policy offline, over distribution of parametric instances in batched setting.
So, what about safety?
Learning Stable DPC Policies with Neural Lyapunov Functions
28
Sayak Mukherjee, Ján Drgoňa, Aaron Tuor, Mahantesh Halappanavar, Draguna Vrabie, Neural Lyapunov Differentiable Predictive Control, Conference on Decision and Control (CDC), 2022
Learning Safe DPC Policies with Control Barrier Functions
29
Wenceslao Shaw Cortez, Ján Drgoňa, Aaron Tuor, Mahantesh Halappanavar, Draguna Vrabie, Differentiable Predictive Control with Safety Guarantees: A Control Barrier Function Approach, Conference on Decision and Control (CDC), 2022
Co-Authors of Differentiable Predictive Control with Neural Lyapunov Functions and Control Barrier Functions
30
Aaron Tuor
Draguna
Vrabie
Ján Drgoňa
Wenceslao Shaw Cortez
Sayak Mukherjee
Mahantesh Halappanavar
Some Open Challenges
Mixed-Integer Decision Space
Modeling of Differential
Algebraic Equations (DAEs)
Control of Partial Differential Equations (PDEs)
D. R. Sharkar, J. Drgoňa, S. Goswami, "Learning to Control PDEs with Differentiable Predictive Control and Time-Integrated Neural Operators," under review, 2025
J. Koch, M. Shapiro, H. Sharma, D. Vrabie, J. Drgoňa, Learning Neural Differential Algebraic Equations via Operator Splitting, CDC, 2025.
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Some Open Challenges
Mixed-Integer Decision Space
Modeling of Differential
Algebraic Equations (DAEs)
Control of Partial Differential Equations (PDEs)
D. R. Sharkar, J. Drgoňa, S. Goswami, "Learning to Control PDEs with Differentiable Predictive Control and Time-Integrated Neural Operators," under review, 2025
J. Koch, M. Shapiro, H. Sharma, D. Vrabie, J. Drgoňa, Learning Neural Differential Algebraic Equations via Operator Splitting, CDC, 2025.
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Session WeC04: �Physics-Aware Learning for Planning and Control
16:30-18:30 �Oceania IV
Some Open Challenges
Mixed-Integer Decision Space
Modeling of Differential
Algebraic Equations (DAEs)
Control of Partial Differential Equations (PDEs)
D. R. Sharkar, J. Drgoňa, S. Goswami, "Learning to Control PDEs with Differentiable Predictive Control and Time-Integrated Neural Operators," under review, 2025
J. Koch, M. Shapiro, H. Sharma, D. Vrabie, J. Drgoňa, Learning Neural Differential Algebraic Equations via Operator Splitting, CDC, 2025.
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
34
D. R. Sharkar, J. Drgoňa, S. Goswami, "Learning to Control PDEs with Differentiable Predictive Control and Time-Integrated Neural Operators," arXiv:2511.08992, 2025
Learning to Control PDEs with DPC and Neural Operators
35
D. R. Sharkar, J. Drgoňa, S. Goswami, "Learning to Control PDEs with Differentiable Predictive Control and Time-Integrated Neural Operators," arXiv:2511.08992, 2025
Time-Integrated Neural Operator (TI-DeepOnet)
⊙ denotes element-wise multiplication
is a state branch net encoding the solution field
is a control branch net encoding the control function
is a trunk net encoding the solution at the collocation points
36
D. R. Sharkar, J. Drgoňa, S. Goswami, "Learning to Control PDEs with Differentiable Predictive Control and Time-Integrated Neural Operators," arXiv:2511.08992, 2025
Formulation of DPC with TI-DeepOnet
37
Learning to Control PDEs with DPC and Neural Operators
D. R. Sharkar, J. Drgoňa, S. Goswami, "Learning to Control PDEs with Differentiable Predictive Control and Time-Integrated Neural Operators," arXiv:2511.08992, 2025
Co-Authors of Learning to Control PDEs with Differentiable Predictive Control and Time-Integrated Neural Operators
38
Dibakar Roy Sharkar
Somdatta Goswami
Ján Drgoňa
Some Open Challenges
Mixed-Integer Decision Space
Modeling of Differential
Algebraic Equations (DAEs)
Control of Partial Differential Equations (PDEs)
D. R. Sharkar, J. Drgoňa, S. Goswami, "Learning to Control PDEs with Differentiable Predictive Control and Time-Integrated Neural Operators," under review, 2025
J. Koch, M. Shapiro, H. Sharma, D. Vrabie, J. Drgoňa, Learning Neural Differential Algebraic Equations via Operator Splitting, CDC, 2025.
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Challenges of Existing Learning to Optimize Methods
40
1. Collecting Solutions as Training Labels for Supervised Learning is Very Expensive
2. Neural Networks Cannot Directly Output Integer Values
Our Solution: Self-Supervised Learning Approach without requiring solutions for training.
Our Solution: Differentiable Integer Correction Layers to ensure integer feasibility.
Our Solution: Gradient-based Feasibility Projection to guarantee feasible integer solutions.
3. It is Difficult to Ensure Feasibility, Especially in Integers
Selected Existing L2O methods�Ferdinando Fioretto, et al., Predicting ac optimal power flows: Combining deep learning and lagrangian dual methods. AAAI conference on AI, 2020.
Priya Donti, et al., DC3: A learning method for optimization with hard constraints, ICLR, 2021
James Kotary, et al., End-to-end constrained optimization learning: A survey. arXiv preprint arXiv:2103.16378, 2021.
He He, et al., Learning to search in branch and bound algorithms. NeurIPS, 2014.
Elias Khalil, et al., Learning to branch in mixed integer programming. AAAI Conference on AI, 2016.
Maxime Gasse, et al., Exact combinatorial optimization with graph convolutional neural networks. NeruIPS, 2019.
Timo Berthold , et al., Learning to scale mixed-integer programs. AAAI Conference on AI, 2021.
Yoshua Bengio, et al., Machine learning for combinatorial optimization: a methodological tour d’horizon. European Journal of Operational Research, 2021.
Dimitris Bertsimas and Bartolomeo Stellato. Online mixed-integer optimization in milliseconds. INFORMS Journal on Computing, 2022.
Learning to Optimize for Mixed-Integer Nonlinear Programming
41
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Learning problem formulation:
Penalty loss function reformulation:
Amortized optimization allows scaling to some of the largest MINLPs.
Could also be used as primal heuristics.
Learning to Optimize for Mixed-Integer Nonlinear Programming
42
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Learning problem formulation:
Penalty loss function reformulation:
Learning to Optimize for Mixed-Integer Nonlinear Programming
43
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Learning problem formulation:
Penalty loss function reformulation:
Learning to Optimize for Mixed-Integer Nonlinear Programming
44
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Learning problem formulation:
Penalty loss function reformulation:
Until now this is standard self-supervised L2O for continuous problems.
Learning to Optimize for Mixed-Integer Nonlinear Programming
45
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Learning problem formulation:
Penalty loss function reformulation:
Main innovation of the paper: Integer correction layers and integer feasibility projection
Differentiable Integer Correction Layers
46
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Learnable end-to-end extension of the Relaxation Enforced Neighborhood Search (RENS).
Differentiable Integer Correction Layers: RC Example
47
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Integer Feasibility Projection
48
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Learning-based alternative to Feasibility Pump which also alternates between rounding and projection.
Approximate Feasibility Guarantees for Integer Projection
49
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
L2O for MINLP: Comparison with SOTA Solvers
50
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Exact solvers such as Gurobi and SCIP can find better solutions over time but are slow. In contrast, our methods achieve high-quality feasible solutions within milliseconds. Our methods provide up to 5 orders of magnitude speedup.
Effect of Penalty Weights and Integer Projections on Feasibility
51
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Our approach achieves comparable or better feasible solutions compared to exact solvers.
Co-Authors of Learning to Optimize for Mixed-Integer Non-Linear Programming
52
Ján Drgoňa
Associate Professor
Department of Civil and Systems
Johns Hopkins University
Bo Tang
PhD Candidate
Department of Mechanical & Industrial Engineering
University of Toronto
Elias B. Khalil
Assistant Professor
Department of Mechanical & Industrial Engineering
University of Toronto
53
Ján Boldocký, Shahriar Dadras Javan, Martin Gulan, Martin Mönnigmann, Ján Drgoňa, "Learning to Solve Parametric Mixed-Integer Optimal Control Problems via Differentiable Predictive Control," arXiv:2506.19646, 2025
Mixed-Integer Differentiable Predictive Control
Co-Authors of Learning to Control PDEs with Differentiable Predictive Control and Time-Integrated Neural Operators
54
Ján Drgoňa
Martin Mönnigmann
Martin Gulan
Shahriar Dadras Javan
Ján Boldocký
Summary
55
NeuroMANCER Scientific Machine Learning Library
56
Open-source library in PyTorch
NeuroMANCER Team
57
Aaron Tuor
Draguna
Vrabie
James Koch
Madelyn Shapiro
Rahul
Birmiwal
Bruno
Jacob
Ján Drgoňa
NeuroMANCER Scientific Machine Learning Library
58
import neuromancer as nm
p = nm.variable(‘p’)
x = nm.variable(‘x’)
y = nm.variable('y’)
�obj = ((1-x)**2 + p*(y-x**2)**2).minimize(weight=1.0, name='obj’)�c1 = (p/2)**2 <= x**2 + y**2
c2 = x**2 + y**2 <= p**2
c3 = x >= y��net = nm.MLP(insize=2, outsize=2, hsizes=[80]*4)�map = nm.Node(net, input_keys=['p’], output_keys=[‘x’,‘y’])
loss = nm.PenaltyLoss([obj], [c1, c2, c3])�problem = nm.Problem([map], loss)
optimizer = torch.optim.AdamW(problem.parameters())�trainer = nm.Trainer(problem,data,optimizer)
best_model = trainer.train()
2. Python code interface
1. Mathematical formulation
4. Results
3. Problem graph
map
Learning to Control Building Energy System
59
J. Drgona, et al., Physics-constrained deep learning of multi-zone building thermal dynamics, Energy and Buildings, 2021
J. Drgona, et al., Deep Learning Explicit Differentiable Predictive Control Laws for Buildings, IFAC NMPC 2021
Benefits of Scientific Machine Learning
Modeling and optimal control design is roughly 10-times faster and requires less expertise.
Real-time decisions are made orders of magnitude faster than traditional model-based approaches.
Learning to Control Power System
60
Ethan King, et al., Koopman-based Differentiable Predictive Control for the Dynamics-Aware Economic Dispatch Problem, American Control Conference 2022
Benefits of Scientific Machine Learning
Fast prototyping by re-using code template from building control project.
Real-time decisions are made orders of magnitude faster than traditional model-based approaches.
Asymptotic Guarantees of Integer Feasibility Projection
61
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Non-Asymptotic Guarantees of Integer Feasibility Projection
62
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Non-Asymptotic Guarantees of Integer Feasibility Projection
63
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Effect of Penalty Weight
64
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Integer feasibility projection (RC-P and LT-P), improves feasibility even with smaller penalty weights. Without projection there is a trade-off between feasibility and objective values (RC and LT).
Benchmark Methods
65
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Method | Description |
EX (Exact Solver) | Solves problems exactly using traditional solver with 1000-sec time-limit as a benchmark. |
N1 (Root Node Solution) | Finds the first feasible solution from the root node of the solver, combining various heuristics. |
RC (Rounding Classification) | A neural network-based correction layer that learns a classification to determine how to round each integer variable. |
LT (Learnable Thresholding) | A neural network-based correction layer that learns a threshold value to decide to round up or down for each integer variable. |
RC-P (RC + Feasibility Projection) | RC combined with feasibility projection, which corrects infeasibilities while preserving integer constraints. |
LT-P (LT + Feasibility Projection) | LT combined with feasibility projection, which corrects infeasibilities while preserving integer constraints. |
L2O for MINLP: Empirical Evaluation
66
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Method | RC | RC-P | LT | |||||||||
Metric | Obj Mean | Obj Median | % Feasible | Time (Sec) | Obj Mean | Obj Median | % Feasible | Time (Sec) | Obj Mean | Obj Median | % Feasible | Time (Sec) |
100×100 | -13.54 | -13.6 | 96% | 0.0022 | -13.54 | -13.57 | 100% | 0.005 | -13.65 | -13.77 | 93% | 0.0023 |
200×200 | -31.62 | -31.71 | 97% | 0.0021 | -31.62 | -31.71 | 100% | 0.005 | -31.34 | -31.61 | 95% | 0.0022 |
500×500 | -73.31 | -73.38 | 86% | 0.0025 | -73.31 | -73.38 | 100% | 0.0065 | -72.36 | -72.48 | 94% | 0.0026 |
1000×1000 | -142.7 | -142.7 | 82% | 0.0042 | -142.7 | -142.7 | 100% | 0.009 | -142.6 | -142.6 | 100% | 0.0047 |
Method | LT-P | EX | N1 | |||||||||
Metric | Obj Mean | Obj Median | % Feasible | Time (Sec) | Obj Mean | Obj Median | % Feasible | Time (Sec) | Obj Mean | Obj Median | % Feasible | Time (Sec) |
100×100 | -13.65 | -13.77 | 100% | 0.01 | -20.79 | -20.78 | 100% | 1237 | 1.5E+18 | 1.4E+18 | 100% | 104.2 |
200×200 | -31.34 | -31.61 | 100% | 0.0064 | - | - | - | - | - | - | - | - |
500×500 | -72.36 | -72.48 | 100% | 0.0063 | - | - | - | - | - | - | - | - |
1000×1000 | -142.6 | -142.6 | 100% | 0.0086 | - | - | - | - | - | - | - | - |
Integer Quadratic Problems (IQPs). Each problem size is evaluated on a test set of 100 instance.
L2O for MINLP: Empirical Evaluation
67
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Integer Non-convex Problems (INPs). Each problem size is evaluated on a test set of 100 instance.
Method | RC | RC-P | LT | |||||||||
Metric | Obj Mean | Obj Median | % Feasible | Time (Sec) | Obj Mean | Obj Median | % Feasible | Time (Sec) | Obj Mean | Obj Median | % Feasible | Time (Sec) |
100×100 | 1.664 | 1.594 | 100% | 0.0022 | 1.664 | 1.594 | 100% | 0.0060 | 0.669 | 0.649 | 96% | 0.0021 |
200×200 | 1.472 | 1.436 | 99% | 0.0022 | 1.471 | 1.436 | 100% | 0.0054 | -0.356 | -0.373 | 100% | 0.0023 |
500×500 | 0.526 | 0.526 | 96% | 0.0029 | 0.524 | 0.526 | 100% | 0.0061 | -1.374 | -1.594 | 98% | 0.0029 |
1000×1000 | 1.423 | 0.809 | 97% | 0.0040 | 1.423 | 0.809 | 100% | 0.0115 | -3.744 | -3.716 | 99% | 0.0050 |
Method | LT-P | EX | N1 | |||||||||
Metric | Obj Mean | Obj Median | % Feasible | Time (Sec) | Obj Mean | Obj Median | % Feasible | Time (Sec) | Obj Mean | Obj Median | % Feasible | Time (Sec) |
100×100 | 0.669 | 0.649 | 100% | 0.0058 | 256.93 | 134.62 | 14% | 1001 | 4411 | 155.2 | 14% | 940.4 |
200×200 | -0.356 | -0.373 | 100% | 0.0056 | - | - | - | - | - | - | - | - |
500×500 | -1.374 | -1.594 | 100% | 0.0072 | - | - | - | - | - | - | - | - |
1000×1000 | -3.744 | -3.716 | 100% | 0.0117 | - | - | - | - | - | - | - | - |
L2O for MINLP: Empirical Evaluation
68
Bo Tang, Elias B. Khalil, Ján Drgoňa, Learning to Optimize for Mixed-Integer Non-linear Programming, arXiv:2410.11061, 2024
Mixed-integer Rosenbrock Problems (MIRBs). Each problem size is evaluated on a test set of 100 instance.
Method | RC | RC-P | LT | |||||||||
Metric | Obj Mean | Obj Median | % Feasible | Time (Sec) | Obj Mean | Obj Median | % Feasible | Time (Sec) | Obj Mean | Obj Median | % Feasible | Time (Sec) |
20×4 | 59.39 | 48.86 | 100% | 0.0019 | 59.39 | 48.86 | 100% | 0.0048 | 62.51 | 63.40 | 100% | 0.0020 |
200×4 | 503.5 | 461.7 | 99% | 0.0021 | 504.2 | 461.7 | 100% | 0.0052 | 622.8 | 626.0 | 100% | 0.0026 |
50×4 | 5938 | 5792 | 99% | 0.0033 | 5942 | 5792 | 100% | 0.0070 | 5612 | 5558 | 97% | 0.0030 |
1000×4 | 6.7E+4 | 6.7E+4 | 76% | 0.0121 | 9.8E+4 | 7.3E+4 | 100% | 0.0824 | 4.8E+4 | 3.5E+4 | 66% | 0.0127 |
Method | LT-P | EX | N1 | |||||||||
Metric | Obj Mean | Obj Median | % Feasible | Time (Sec) | Obj Mean | Obj Median | % Feasible | Time (Sec) | Obj Mean | Obj Median | % Feasible | Time (Sec) |
20×4 | 62.51 | 63.40 | 100% | 0.0055 | 64.67 | 59.16 | 100% | 1005 | 87.83 | 77.34 | 100% | 0.0813 |
200×4 | 622.8 | 626.0 | 100% | 0.0062 | 8.4E+5 | 908.8 | 100% | 1002 | 3.7E+8 | 957.4 | 100% | 0.2608 |
2000×4 | 5615 | 5558 | 100% | 0.0071 | 4.7E+10 | 9262 | 96% | 1002 | 8.3E+12 | 9379 | 95% | 71.91 |
20000×4 | 8.0E+4 | 4.5E+4 | 100% | 0.0639 | 1.1E+15 | 1.0E5 | 78% | 1040 | 1.2E+15 | 1.0E5 | 78% | 782.1 |
Metric Learning to Accelerate Convergence of Operator Splitting Methods
69
Douglas-Rachford splitting (DR) algorithm:
Parametric programming setting:
Ethan King, James Kotary, Ferdinando Fioretto, Jan Drgona, Metric Learning to Accelerate Convergence of Operator Splitting Methods for Differentiable Parametric Programming, Under review for CDC 2024.
Idea: Train neural network to optimize the metric as a function of problem parameters:
We can accelerate convergence of DR and ADMM algorithms via end-to-end metric learning.
Metric Learning is a Form of Active Set Prediction
70
Ethan King, James Kotary, Ferdinando Fioretto, Jan Drgona, Metric Learning to Accelerate Convergence of Operator Splitting Methods for Differentiable Parametric Programming, Under review for CDC 2024.
TODO: change this slide to be motivating and more flashy
71
J. Koch, M. Shapiro, H. Sharma, D. Vrabie, J. Drgona, Learning Neural Differential Algebraic Equations via Operator Splitting, arXiv:2403.12938, 2024.
DAE parameter estimation problem:
The Picard–Lindelöf theorem (also called the Cauchy–Lipschitz theorem) gives sufficient conditions under which an initial value problem (IVP) for an ordinary differential equation (ODE) has a unique solution.
Assumptions:
[30] K. E. Brenan, S. L. Campbell, and L. R. Petzold, Numerical Solution of Initial-Value Problems in Differential-Algebraic Equations. SIAM, 1996.
Networked Dynamical Systems via Universal Differential Equations (UDEs)
Network of unknown oscillators
Coupling adjacency
Ground truth Kuramoto system:
Generalization of learned dynamics on never-seen network topology.
James Koch, et al., Structural Inference of Networked Dynamical Systems with Universal Differential Equations, Chaos: An interdisciplinary Journal of Nonlinear Science, doi.org/10.1063/5.0109093, 2023