1
Domain Model Learning in AI Planning
Tutorial in AAAI 2026
Part 5: Beyond Offline Learning of Classical Planning Models
Roni Stern
Ben Gurion University of the Negev
Common Setting for Learning Planning Domains
3
Beyond the Common Problem Setting
4
Numeric Planning
5
Action: Move(Truck, A, B)
Pre: At(Truck, A)
Eff: At(Truck, B), not(Truck, A)
State:
At(Truck, A)
At(Package, B)
State:
At(Truck, A)
At(Package, B)
Fuel(Truck)=5
Action: Move(Truck, A, B)
Pre: At(Truck, A),
Fuel(Truck)>0
Eff: At(Truck, B), not(Truck, A)
Decrease Fuel(Truck) by 1
Numeric state variables
Numeric preconditions
Numeric effects
PlanMiner [Segura-Muros et al., 2021]
6
A
B
C
fuel = 2.5
Move
(A, B)
A
B
C
fuel = 1.5
A
B
C
fuel = 0.5
Dataset for Move(?x, ?y)
| Pre | Post | ||||
| Fuel | At(?x) | At(?y) | Fuel | At(?x) | At(?y) |
1 | 2.5 | True | False | 1.5 | False | True |
2 | 1.5 | True | False | 0.5 | False | True |
Load
(Pkg, C)
Move
(B, C)
A
B
C
fuel = 0.5
| Pre | Post | ||||
| Fuel | In(?x) | … | Fuel | In(?x) | … |
1 | 0.5 | False | … | 0.5 | In | … |
Dataset for Load(?x, ?y)
PlanMiner [Segura-Muros et al., 2021]
7
A
B
C
fuel = 2.5
Move
(A, B)
A
B
C
fuel = 1.5
A
B
C
fuel = 0.5
Dataset for Move(?x, ?y)
| Pre | Post | | ||||
| Fuel | At(?x) | At(?y) | Fuel | At(?x) | At(?y) | |
1 | 2.5 | True | False | 1.5 | False | True | 1 |
2 | 1.5 | True | False | 0.5 | False | True | 1 |
Load
(Pkg, C)
Move
(B, C)
A
B
C
fuel = 0.5
| Pre | Post | ||||
| fuel | At(?x,?y) | In(?x,?z) | fuel | At(?x,?y) | At(?x,?y) |
1 | 0.5 | True | False | 0.5 | False | True |
Symbolic Regression over functions and integers
Objective: fit numeric expressions to the post-values
Example: Fuel-1
PlanMiner [Segura-Muros et al., 2021]
8
Dataset for Move(?x, ?y)
| Pre | Post | ||||||
| Fuel | Fuel-1 | At(?x) | At(?y) | Fuel | Fuel-1 | At(?x) | At(?y) |
1 | 2.5 | 1.5 | True | False | 1.5 | 0.5 | False | True |
2 | 1.5 | 0.5 | True | False | 0.5 | -0.5 | False | True |
| Pre | Post | ||||
| fuel | At(?x,?y) | In(?x,?z) | fuel | At(?x,?y) | At(?x,?y) |
1 | 0.5 | True | False | 0.5 | False | True |
Learn Rules to Classify Between Pre- and Post- States (NSLV)
Preconditions!
PlanMiner [Segura-Muros et al., 2021]
9
Advantages:
Limitations:
PlanMiner [Segura-Muros et al., 2021]
10
Learning numeric preconditions
= Learning Boolean functions
Learning numeric effects
= Learning general functions
Intractable to learn ☹
Solution: reasonable assumption ☺
Effects are
linear functions
Preconditions are
conjunctions of linear inequality
Numeric SAM (N-SAM) [Mordoch et al. ’23]
11
N-SAM
Numeric
Discrete
SAM Learning
Learning half-spaces
Linear regression
Preconditions
Effects
Numeric SAM: Learning Effects
12
Pre-state vectors as matrix
Coefficient vector
Next state values
Safety:
Only allow action with a unique fit
Pre-state
Action
Post-state
Numeric SAM: Learning Preconditions
13
13
Observation #1:
Preconditions form a convex hull
Numeric SAM: Learning Preconditions
14
14
Observation #1:
Preconditions form a convex hull
Observation #2:
Sampled pre-states form a “sub” convex hull
Numeric SAM: Learning Preconditions
15
15
Observation #1:
Preconditions form a convex hull
Observation #2:
Sampled pre-states form a “sub” convex hull
Limitation: Linear Dependencies
16
Cannot create a convex hull!
NSAM* [Mordoch et al. ’24]
Technical details:
17
Cannot create a convex hull!
Project points to the
subspace spanned by them
NSAM* is optimal for learning safe action models
(optimal=cannot learn stronger action model with the same data)
Limitations: Sample Complexity
18
Reachable states
?
Sampled states
New state
Learned Convex Hull
Experimental Results for N-SAM
19
Domain | Numeric Recall | MSE | Discrete Precision | Discrete Recall |
Farmland | 1.00 | 0.00 | 0.50 | 1.00 |
Driverlog | 1.00 | 0.00 | 0.64 | 1.00 |
Depots | 0.99 | 0.00 | 0.77 | 1.00 |
Sailing | 0.99 | 0.00 | 1.00 | 1.00 |
Counters | 0.74 | 0.00 | - | - |
Satellite | 0.99 | 0.00 | 0.74 | 1.00 |
Rovers | 0.94 | 0.00 | 0.58 | 0.84 |
Zenotravel | 0.99 | 0.00 | 0.84 | 1.00 |
Benchmark Numeric Planning Problems from the IPC and Scala et al.
Problem Solving Capabilities
20
Note: PlanMiner is not safe
Note: only 80 trajectories!
Learning Numeric Planning Domain
N-PlanMiner
☺ Handles arbitrary numerical expressions (*)
☺ Some support for noise and missing values
☹ No theoretical guarantees, runtime is unbounded
NSAM*
☺ Runtime is polynomial in the input transition
Many open directions:
….
21
That’s it?!
(Juba and Stern ’22)
(Deng and Juba ‘23)
Beyond the Common Problem Setting
22
Multi-Agent
23
From One to Many [Brafman and Domshlak ‘08]
Step | 1 | 3 | 4 | 5 | 6 | 7 | 8 | 9 |
Agent | 1 | 1 | 1 | 2 | 2 | 3 | 3 | 1 |
Action | A11 | A12 | A13 | A21 | A22 | A31 | A33 | A14 |
Lammas [Zhuo et al. ‘11],
MA-SAM [Mordoch et al. ’22]
From One to Many [Brafman and Domshlak ‘08]
Step | 1 | 2 | 3 | 4 | 5 | 6 |
Agent 1 | A11 | | A12 | A13 | | |
Agent 2 | A21 | A22 | | A23 | A24 | A25 |
Agent 3 | | A31 | A32 | | A33 | A34 |
Conflicts between actions?
Collaborative actions?
“Credit assignment”?
Concurrent execution – trajectories of joint actions
From One to Many [Brafman and Domshlak ‘08]
Step | 1 | 2 | 3 | 4 | 5 | 6 |
Agent 1 | A11 | | A12 | A13 | | |
Agent 2 | A21 | A22 | | A23 | A24 | A25 |
Agent 3 | | A31 | A32 | | A33 | A34 |
“Credit assignment”
Concurrent execution – trajectories of joint actions
Not(Have(apple))
Have(apple)
Have(apple) is an effect of
A11, or A21, or both
From One to Many [Brafman and Domshlak ‘08]
Step | 1 | 2 | 3 | 4 | 5 | 6 |
Agent 1 | A11 | | A12 | A13 | | |
Agent 2 | A21 | A22 | | A23 | A24 | A25 |
Agent 3 | | A31 | A32 | | A33 | A34 |
“Credit assignment”
Concurrent execution – trajectories of joint actions
From One to Many [Brafman and Domshlak ‘08]
Step | 1 | 2 | 3 | 4 | 5 | 6 |
Agent 1 | A11 | | A12 | A13 | | |
Agent 2 | A21 | A22 | | A23 | A24 | A25 |
Agent 3 | | A31 | A32 | | A33 | A34 |
“Credit assignment”
Concurrent execution – trajectories of joint actions
Many open directions:
….
Beyond the Common Problem Setting
29
Online Domain Model Learning
30
Action
State
Learning Agent
Online
Learning
Active
Learning
Goal-Literal Babbling (Chitnis et al. ‘21)
31
Action
State
Learning Agent
Online Learning of Action Models [Lamanna et al. ‘21]
32
Action
State
Learning Agent
Online Learning of Action Models [Lamanna et al. ‘21]
33
Online Learning of Action Models [Lamanna et al. ‘21]
34
What is not informative?
and we won’t learn anything new
Online Learning of Action Models [Lamanna et al. ‘21]
35
Action
State
Learning Agent
Halt when no such (s,op) pair exists
Online Learning of Action Models [Lamanna et al. ‘21]
☺ Converge to a sufficient domain model
36
Many open directions:
….
Agent Interrogation Algorithm (AIA) [Verma et al. ‘21]
37
Extended to support user-AI vocabulary differences (Verma et al. ‘22)
Optimistic Exploration (Sreedharan & Katz ‘23)
38
Optimistic Exploration (Sreedharan & Katz ‘23)
39
Used train an off-policy
Model-free RL algorithm
(Q Learning)
Numeric Online Action Model Learning [Mordoch et al. ‘26]
40
Numeric Online Action Model Learning [Mordoch et al. ‘26]
41
Many open directions:
….
Summary and What’s Next?
Many, many direction for future work
42
Online representation and action model learning
Richer Domains
(H)RL-integrations
(some work on this)
LLM-integrations (some work on this)
Theoretical foundations
Applications: is it really worth it?
What’s Next – Lunch!
43
Time | Session | Speaker |
08:30–09:15 | Introduction & Domain Learning Basics | Roni Stern |
09:15–09:45 | Learning State Abstractions | Roni Stern |
09:45–10:30 | Offline Learning Domain Models | Leonardo Lamanna |
10:30–11:00 | Coffee Break | |
11:00–11:45 | Hands-on Session | Leonardo Lamanna |
11:45–12:30 | Online Learning and Open Challenges | Roni Stern |
Thank you!
Link to feedback form:
please help us get better