Dr. Khellat Kihel Souad
Machine Learning (ML)
REPUBLIQUE ALGERIENNE DEMOCRATIQUE ET POPULAIRE
Ministère de l'Enseignement Supérieur et de la Recherche Scientifique
Université des sciences et de la technologie Mohamed-Boudiaf, Oran.
Faculté des mathématiques et informatique
Département d'informatique
2
CHAPTER 5 (PART II) : REINFORCEMENT LEARNING�
3
PROBLEMATICS
4
5
6
7
INTRODUCTION TO REINFORCEMENT LEARNING
8
A MULTI-DICIPLINARY FIELD
9
INTRODUCTION
How to learn to navigate a maze without making mistakes?�Or, how to master a good strategy for:
10
REINFORCEMENT LEARNING
11
REINFORCEMENT LEARNING
12
GENERAL CASE
An agent operates in a given environment.�It can perform certain actions based on the current state:
The state of the environment, And its own internal state,�Leading to a transition to a new state.
13
Reinforcement Learning:�A framework for adapting an agent to its environment by leveraging rewards/punishments (reinforcement signals).
14
KEY POINTS
Reinforcement Learning (RL) involves several fundamental concepts that define how machines learn from experience and make decisions:
15
THE AGENT-ENVIRONMENT INTERACTION PROTOCOL
16
REINFORCEMENT LEARNING PROCESS
17
ENVIRONMENT
18
THE CRITIC & AGENT
19
MARKOV DECISION PROCESS
A Markov Decision Process (MDP) is a mathematical framework that provides a structured way to model environments in reinforcement learning.
An MDP is formally defined by the tuple (S, A, T, R, γ), where:
20
21
BELLMAN EQUATION
The Bellman equation computes the value of being in a state or taking an action based on expected future rewards.
It decomposes the total expected reward into:
22
23
Q-LEARNING ALGORITHM
Q-Learning is a model-free algorithm, meaning it doesn’t require prior knowledge of the environment’s dynamics (transition probabilities or reward structure). Instead, it learns by direct interaction with the environment.
24
KEY CONCEPT OF Q-LEARNING
1. Q-Value (Q(s,a))
2. Q-Table
Structure: A lookup table storing Q-values for all possible (state, action) pairs.
Learning: Updated iteratively as the agent explores the environment.
25
State (s) | Action (a) | Q(s,a) |
s₁ | a₁ | 0.5 |
s₁ | a₂ | 1.2 |
s₂ | a₁ | -0.3 |
26
3. Learning Rate (α, 0 ≤ α ≤ 1)
4. Discount Factor (γ, 0 ≤ γ < 1)
Why These Matter
27
28
29
Q LEARNING PROCESS
The Q-learning algorithm does not specify a fixed number of updates for the Q-table. Instead:
Key Idea:
30
SARSA
31
32
Aspect | Q-Learning | SARSA |
Type | Off-policy | On-policy |
Action utilisée pour maj | Meilleure action possible | Action réellement choisie |
Risque | Peut apprendre des politiques risquées | Apprend des politiques plus prudentes |
Apprentissage | Plus rapide mais parfois moins sûr | Plus lent mais plus stable |
Ex : Agent sur une falaise (Cliff Walking)