1 of 9

Hackathon: Reinforcement Learning

OpenAI Gym

Team: Radio Frequency

Team members:

Jarek Nachyla

Tommaso Van Der Meer

Konstantinos Marios Georgantas

Antonio Mario Pio Frioli

Gerard Tomczynski

2 of 9

Reinforcement learning challenges:

  • Lunar Lander
  • Cartpole
  • Lunar Lander Video
  • Cartpole Video
  • Bipedal walker Video

3 of 9

Cart pole - goal

  • Cartpole can’t excess 15 degrees from vertical position
  • Cartpole can’t move more than 2.4 unit from center
  • Keep the balance of the cart pole as long as possible

4 of 9

Cart pole - deep Q learning network

  • we build our own DQN model in tensorflow framework
  • Neural Network that inputs the actions and outputs the Q-Table values
  • The agent will pick the max Q Value for that state

5 of 9

Cart pole - algorithm, parameters

  • Replays (sampling from history to avoid network forgetting )
  • Epsilon greedy algorithm with decay
  • Learning Rate = 0.0001, gamma = 0.95

6 of 9

Cart pole - results

  • 100 epochs each 1000 episodes
  • X axis epochs
  • Y axis - score (sum of rewards)

7 of 9

Lunar lander - goal

  • Land the spaceship between the flags (landing pad)
  • Spaceship can’t land outside of landing pad
  • Land with the low velocity
  • Use as little fuel as possible

8 of 9

Lunar lander - Algorithm

  • Proximal Policy Optimization algorithm from stable_baseline library
  • MLP policy
  • Maximizes the reward from the beginning differently from Q-Table based algorithms

9 of 9

Lunar lander - Training

  • number training steps 100k
  • we use SaveBestTrainigCallback

in order to save best model