1 of 43

1

Domain Model Learning in AI Planning

Tutorial in AAAI 2026

2 of 43

Part 5: Beyond Offline Learning of Classical Planning Models

Roni Stern

Ben Gurion University of the Negev

3 of 43

Common Setting for Learning Planning Domains

  • States include Boolean state variables

  • Deterministic effects

  • Single agent

  • Offline learning

3

4 of 43

Beyond the Common Problem Setting

  • States include Boolean state variables
    • States also include numeric state variables
  • Deterministic effects
    • Stochastic effects, partial observability
  • Single agent
    • Multiple agents acting concurrenlty
  • Offline learning
    • Active learning: must act to collect observations

4

5 of 43

Numeric Planning

5

Action: Move(Truck, A, B)

Pre: At(Truck, A)

Eff: At(Truck, B), not(Truck, A)

State:

At(Truck, A)

At(Package, B)

State:

At(Truck, A)

At(Package, B)

Fuel(Truck)=5

Action: Move(Truck, A, B)

Pre: At(Truck, A),

Fuel(Truck)>0

Eff: At(Truck, B), not(Truck, A)

Decrease Fuel(Truck) by 1

Numeric state variables

Numeric preconditions

Numeric effects

6 of 43

PlanMiner [Segura-Muros et al., 2021]

  1. Pre-process to extract pre- and post-state datasets per action
  2. Search for numerical expressions that fit the effects
  3. Search for logical rules over them that fit as preconditions

6

  • Discovering relational and numerical expressions from plan traces for learning action models. Segura-Muros, José Á., Juan Fernández-Olivares, and Raúl Pérez. Applied Intelligence, 2021.
  • Learning Numerical Action Models from Noisy Input Data. Segura-Muros, José Á., Juan Fernández-Olivares, and Raúl Pérez. Arxiv, 2021.

A

B

C

fuel = 2.5

Move

(A, B)

A

B

C

fuel = 1.5

A

B

C

fuel = 0.5

Dataset for Move(?x, ?y)

Pre

Post

Fuel

At(?x)

At(?y)

Fuel

At(?x)

At(?y)

1

2.5

True

False

1.5

False

True

2

1.5

True

False

0.5

False

True

Load

(Pkg, C)

Move

(B, C)

A

B

C

fuel = 0.5

Pre

Post

Fuel

In(?x)

Fuel

In(?x)

1

0.5

False

0.5

In

Dataset for Load(?x, ?y)

7 of 43

PlanMiner [Segura-Muros et al., 2021]

  1. Pre-process to extract pre- and post-state datasets per action
  2. Search for numerical expressions that fit the effects
  3. Search for logical rules over them that fit as preconditions

7

  • Discovering relational and numerical expressions from plan traces for learning action models. Segura-Muros, José Á., Juan Fernández-Olivares, and Raúl Pérez. Applied Intelligence, 2021.
  • Learning Numerical Action Models from Noisy Input Data. Segura-Muros, José Á., Juan Fernández-Olivares, and Raúl Pérez. Arxiv, 2021.

A

B

C

fuel = 2.5

Move

(A, B)

A

B

C

fuel = 1.5

A

B

C

fuel = 0.5

Dataset for Move(?x, ?y)

Pre

Post

Fuel

At(?x)

At(?y)

Fuel

At(?x)

At(?y)

1

2.5

True

False

1.5

False

True

1

2

1.5

True

False

0.5

False

True

1

Load

(Pkg, C)

Move

(B, C)

A

B

C

fuel = 0.5

Pre

Post

fuel

At(?x,?y)

In(?x,?z)

fuel

At(?x,?y)

At(?x,?y)

1

0.5

True

False

0.5

False

True

Symbolic Regression over functions and integers

Objective: fit numeric expressions to the post-values

Example: Fuel-1

8 of 43

PlanMiner [Segura-Muros et al., 2021]

  1. Pre-process to extract pre- and post-state datasets per action
  2. Search for numerical expressions that fit the effects
  3. Search for logical rules over them that fit as preconditions

8

  • Discovering relational and numerical expressions from plan traces for learning action models. Segura-Muros, José Á., Juan Fernández-Olivares, and Raúl Pérez. Applied Intelligence, 2021.
  • Learning Numerical Action Models from Noisy Input Data. Segura-Muros, José Á., Juan Fernández-Olivares, and Raúl Pérez. Arxiv, 2021.

Dataset for Move(?x, ?y)

Pre

Post

Fuel

Fuel-1

At(?x)

At(?y)

Fuel

Fuel-1

At(?x)

At(?y)

1

2.5

1.5

True

False

1.5

0.5

False

True

2

1.5

0.5

True

False

0.5

-0.5

False

True

Pre

Post

fuel

At(?x,?y)

In(?x,?z)

fuel

At(?x,?y)

At(?x,?y)

1

0.5

True

False

0.5

False

True

Learn Rules to Classify Between Pre- and Post- States (NSLV)

Preconditions!

9 of 43

PlanMiner [Segura-Muros et al., 2021]

  1. Pre-process to extract pre- and post-state datasets per action
  2. Search for numerical expressions that fit the effects
  3. Search for logical rules over them that fit as preconditions

9

  • Discovering relational and numerical expressions from plan traces for learning action models. Segura-Muros, José Á., Juan Fernández-Olivares, and Raúl Pérez. Applied Intelligence, 2021.
  • Learning Numerical Action Models from Noisy Input Data. Segura-Muros, José Á., Juan Fernández-Olivares, and Raúl Pérez. Arxiv, 2021.

Advantages:

  • Can handle diverse numeric expression
  • Include some handling of missing values
  • N-PlanMiner: extended to handle noisy observations
  • Can incorporate domain knowledge

Limitations:

  • Runtime is unbounded (symbolic regression…)
  • Unclear how to constant numbers
  • No guarantees on soundness or completeness

10 of 43

PlanMiner [Segura-Muros et al., 2021]

  1. Pre-process to extract pre- and post-state datasets per action
  2. Search for numerical expressions that fit the effects
  3. Search for logical rules over them that fit as preconditions

10

  • Discovering relational and numerical expressions from plan traces for learning action models. Segura-Muros, José Á., Juan Fernández-Olivares, and Raúl Pérez. Applied Intelligence, 2021.
  • Learning Numerical Action Models from Noisy Input Data. Segura-Muros, José Á., Juan Fernández-Olivares, and Raúl Pérez. Arxiv, 2021.

Learning numeric preconditions

= Learning Boolean functions

Learning numeric effects

= Learning general functions

Intractable to learn ☹

Solution: reasonable assumption ☺

Effects are

linear functions

Preconditions are

conjunctions of linear inequality

11 of 43

Numeric SAM (N-SAM) [Mordoch et al. ’23]

11

N-SAM

Numeric

Discrete

SAM Learning

Learning half-spaces

Linear regression

Preconditions

Effects

12 of 43

Numeric SAM: Learning Effects

12

 

 

 

Pre-state vectors as matrix

Coefficient vector

Next state values

Safety:

Only allow action with a unique fit

 

Pre-state

Action

Post-state

13 of 43

Numeric SAM: Learning Preconditions

13

13

Observation #1:

Preconditions form a convex hull

14 of 43

Numeric SAM: Learning Preconditions

14

14

Observation #1:

Preconditions form a convex hull

Observation #2:

Sampled pre-states form a “sub” convex hull

15 of 43

Numeric SAM: Learning Preconditions

15

15

Observation #1:

Preconditions form a convex hull

Observation #2:

Sampled pre-states form a “sub” convex hull

16 of 43

Limitation: Linear Dependencies

16

Cannot create a convex hull!

17 of 43

NSAM* [Mordoch et al. ’24]

Technical details:

  • Ensure only states on this subspace are allowed
    • Find basis of orthogonal half-space
    • Preconditions are dot product = 0
  • Preconditions need to be projected as well

17

Cannot create a convex hull!

Project points to the

subspace spanned by them

NSAM* is optimal for learning safe action models

(optimal=cannot learn stronger action model with the same data)

18 of 43

Limitations: Sample Complexity

  • Worst case is linear in the # states ☹
  • Open question: does this occurs “in practice”?

18

Reachable states

?

Sampled states

New state

Learned Convex Hull

19 of 43

Experimental Results for N-SAM

19

Domain

Numeric Recall

MSE

Discrete Precision

Discrete Recall

Farmland

1.00

0.00

0.50

1.00

Driverlog

1.00

0.00

0.64

1.00

Depots

0.99

0.00

0.77

1.00

Sailing

0.99

0.00

1.00

1.00

Counters

0.74

0.00

-

-

Satellite

0.99

0.00

0.74

1.00

Rovers

0.94

0.00

0.58

0.84

Zenotravel

0.99

0.00

0.84

1.00

Benchmark Numeric Planning Problems from the IPC and Scala et al.

20 of 43

Problem Solving Capabilities

20

  1. Better than PlanMiner in 4 (sometimes drastically)
  2. Often NSAM is sufficient and NSAM* is not needed

  • PlanMiner better in 3 domains

  • Inconclusive results in 2 domains

Note: PlanMiner is not safe

Note: only 80 trajectories!

21 of 43

Learning Numeric Planning Domain

N-PlanMiner

☺ Handles arbitrary numerical expressions (*)

☺ Some support for noise and missing values

☹ No theoretical guarantees, runtime is unbounded

NSAM*

☺ Runtime is polynomial in the input transition

  • Guaranteed safety and optimality (sample complexity)
  • Limited to polynomial effects and preconditions
  • No support for missing observations or noise

Many open directions:

  • Richer domains: stochastic effects, processes and events,…
  • Noise and missing values with guarantees on learned model
  • Integration with RL and neural-network-based methods

….

21

That’s it?!

(Juba and Stern ’22)

(Deng and Juba ‘23)

22 of 43

Beyond the Common Problem Setting

  • States include Boolean state variables
    • States also include numeric state variables
  • Deterministic effects and full observability
    • Stochastic effects, partial observability
  • Single agent
    • Multiple agents acting concurrently
  • Offline learning
    • Active learning: must act to collect observations

22

23 of 43

Multi-Agent

  • The “credit assignment problem” for effects
  • MA-SAM: learn to disambiguate
  • Open question: action interaction, MARL

23

24 of 43

From One to Many [Brafman and Domshlak ‘08]

Step

1

3

4

5

6

7

8

9

Agent

1

1

1

2

2

3

3

1

Action

A11

A12

A13

A21

A22

A31

A33

A14

 

Lammas [Zhuo et al. ‘11],

MA-SAM [Mordoch et al. ’22]

25 of 43

From One to Many [Brafman and Domshlak ‘08]

Step

1

2

3

4

5

6

Agent 1

A11

A12

A13

Agent 2

A21

A22

A23

A24

A25

Agent 3

A31

A32

A33

A34

Conflicts between actions?

Collaborative actions?

“Credit assignment”?

Concurrent execution – trajectories of joint actions

26 of 43

From One to Many [Brafman and Domshlak ‘08]

Step

1

2

3

4

5

6

Agent 1

A11

A12

A13

Agent 2

A21

A22

A23

A24

A25

Agent 3

A31

A32

A33

A34

“Credit assignment”

Concurrent execution – trajectories of joint actions

Not(Have(apple))

Have(apple)

Have(apple) is an effect of

A11, or A21, or both

27 of 43

From One to Many [Brafman and Domshlak ‘08]

Step

1

2

3

4

5

6

Agent 1

A11

A12

A13

Agent 2

A21

A22

A23

A24

A25

Agent 3

A31

A32

A33

A34

“Credit assignment”

Concurrent execution – trajectories of joint actions

 

28 of 43

From One to Many [Brafman and Domshlak ‘08]

Step

1

2

3

4

5

6

Agent 1

A11

A12

A13

Agent 2

A21

A22

A23

A24

A25

Agent 3

A31

A32

A33

A34

“Credit assignment”

Concurrent execution – trajectories of joint actions

Many open directions:

  • Richer domains: stochastic effects, processes and events,…
  • Handle noise and missing values
  • Integration with MARL and neural-network-based methods
  • Learning conflicts between actions and collaborative
  • Distributed learning

….

29 of 43

Beyond the Common Problem Setting

  • States include Boolean state variables
    • States also include numeric state variables
  • Deterministic effects and full observability
    • Stochastic effects, partial observability
  • Single agent
    • Multiple agents acting concurrently
  • Offline learning
    • Active learning: must act to collect observations

29

30 of 43

Online Domain Model Learning

30

Action

State

  • Choose actions
  • Process observations
  • Learn/update domain model

Learning Agent

Online

Learning

Active

Learning

 

31 of 43

Goal-Literal Babbling (Chitnis et al. ‘21)

31

Action

State

  • Choose actions
  • Process observations
  • Learn/update domain model

Learning Agent

 

  • Prioritize novel “goals”
  • Filter impossible goals

32 of 43

Online Learning of Action Models [Lamanna et al. ‘21]

32

Action

State

  • Choose actions
  • Process observations
  • Learn/update domain model

Learning Agent

 

 

33 of 43

Online Learning of Action Models [Lamanna et al. ‘21]

33

 

 

 

 

 

 

34 of 43

Online Learning of Action Models [Lamanna et al. ‘21]

34

 

What is not informative?

 

 

 

 

and we won’t learn anything new

35 of 43

Online Learning of Action Models [Lamanna et al. ‘21]

35

Action

State

  • Choose actions
  • Process observations
  • Learn/update domain model

Learning Agent

 

 

Halt when no such (s,op) pair exists

36 of 43

Online Learning of Action Models [Lamanna et al. ‘21]

☺ Converge to a sufficient domain model

    • Sufficient = safe and complete
  • Requires very few iterations to halt
  • Utilizes failed actions

36

Many open directions:

  • Optimize for limited budget
  • Choice of returned model
  • Richer domains
  • Handle noise and missing values

….

37 of 43

Agent Interrogation Algorithm (AIA) [Verma et al. ‘21]

  • Objective: understand the AI agent
  • Approach: learn its capabilities
    • Iterate over “pal-tuples” (p=predicate, a=action, l=pre. or eff.)
    • Generate a disambiguating plan (to check if pal-tuple is true)
      • Reduce this problem to a planning problem
    • Filter candidate models as needed

37

Extended to support user-AI vocabulary differences (Verma et al. ‘22)

38 of 43

Optimistic Exploration (Sreedharan & Katz ‘23)

  • Maintain an optimistic model
  • Goal-oriented exploration
  • Encourage exploration by using a diverse planner

38

39 of 43

Optimistic Exploration (Sreedharan & Katz ‘23)

  • Maintain an optimistic model
  • Goal-oriented exploration
  • Encourage exploration by using a diverse planner

39

Used train an off-policy

Model-free RL algorithm

(Q Learning)

40 of 43

Numeric Online Action Model Learning [Mordoch et al. ‘26]

  • NOAM: Online learning of a numeric action model
  • Maintains safe and optimistic action models
  • Goal-oriented exploration

40

41 of 43

Numeric Online Action Model Learning [Mordoch et al. ‘26]

  • NOAM: Online learning of a numeric action model
  • Maintains safe and optimistic action models
  • Goal-oriented exploration

41

Many open directions:

  • Scaling to larger domains
  • Optimize for cumulative regret
  • Richer domains
  • Handle noise and missing values

….

42 of 43

Summary and What’s Next?

  • Beyond learning classical planning domains - not much work …
    • Mostly PlanMiner, NSAM, MA-SAM,learning RDDL models…
  • Some work on online learning of domains
    • OLAM, AIA, Optimistic Exploration, Glib, NOAM,…
    • Common approach:
      • Learn an optimistic model
      • Plan to do more informative actions

Many, many direction for future work

42

Online representation and action model learning

Richer Domains

(H)RL-integrations

(some work on this)

LLM-integrations (some work on this)

Theoretical foundations

Applications: is it really worth it?

43 of 43

What’s Next – Lunch!

43

Time

Session

Speaker

08:30–09:15

Introduction & Domain Learning Basics

Roni Stern

09:15–09:45

Learning State Abstractions

Roni Stern

09:45–10:30

Offline Learning Domain Models

Leonardo Lamanna

10:30–11:00

Coffee Break

11:00–11:45

Hands-on Session

Leonardo Lamanna

11:45–12:30

Online Learning and Open Challenges

Roni Stern

Thank you!

Link to feedback form:

please help us get better