1 of 15

SOFT ACTOR-CRITIC SOLUTION TO A SECURITY GAME WITH DECEPTION AND AN INFORMANT

Jinhong K. Guo, Martin O. Hofmann, Carter Veldhuizen, Sidharth Satya, Valerie Champagne

October, 2021

LOCKHEED MARTIN PROPRIETARY INFORMATION

1

2 of 15

OVERVIEW

DECEPTION IS A CRITICAL ENABLER TO WINNING IN PEER CONFLICTS

  • Challenges
  • Technical Approaches
  • Experiments
  • Conclusions

2

3 of 15

CHALLENGES

MAINTAINING DECEPTION OVER EXTENDED ENGAGEMENTS REQUIRES AUTOMATED DECISION SUPPORT

  • Operational Challenges
    • Facing peer adversaries, our military needs additional advantages to create overmatch, to get inside adversaries’ OODA loops by imposing greater complexity, to inflict multiple dilemmas, and to create cognitive overload.
    • Deception promises to delay the adversary’s response, delay recognition of the true course of action, and deter or blunt an attack.
    • Developing deceptive Courses of Action is manpower intensive and time consuming
  • Technical Challenges
    • Complexity of the problem makes a mathematical solution infeasible
    • Learning stability with multiple agents
    • Sparse rewards

3

4 of 15

OUR APPROACH

DECEPTION BECOMES A NATURAL PART OF THE STRATEGY

  • Model competition between two opponents (possibly multiple agents for each) as a two-player plus third-party player, extensive form game.
  • Embed deception in the game theoretic framework by extending the game formulation
  • Apply a centralized learning with decentralized execution algorithm – faster convergence to equilibrium
  • Apply a meta strategy to increase learning efficiency and mitigate reward sparseness
    • Design an abstract game which is solvable with Mixed Integer Linear Programming
    • Use the resulting high level strategy to guide the reinforcement learning algorithm

4

5 of 15

GAME ENVIRONMENT

  • A geographical area with continuous action space
  • A number of targets with known locations
  • Each target has its value to the red player
  • Blue moves at an even speed
  • Red moves at an even speed
  • White moves at an even speed
  • White player can be neutral or an informant for the blue player
  • Red player reduces detection probability by blending in with the neutral player

Neutral White Player

Blue State and Observation:

  • Blue position [x, y]
  • Red position [x, y] when within range of sight
  • White position [x, y]

Red State and Observation:

  • Red position [x, y]
  • Blue position [x, y] when within range of sight
  • White position [x, y]
  • Target location [x, y]

White Player as Informant

Additional Blue State and Observation:

  • White reported red position [x, y] when its near red

5

6 of 15

EXAMPLE STRATEGIES

Red learned to try to blend in with White, but sees blue nearby and starts to pull away from blue.

These snapshots are arranged clockwise starting from the top left chart. Red starts near the lower left corner while blue starts in the middle.

Attempted Deception

6

7 of 15

EXAMPLE STRATEGIES

:

Blue learned to capture red around where it most likely to go with the guidance of the meta strategy

Blue and Red player wait each other out around a target

Foiled Deception

Stalemate

7

8 of 15

STRATEGY COMPUTATION

 

8

9 of 15

EXPERIMENTS

A

Red and blue player cannot see each other; red does not attempt to deceive blue by blending in with white.

B

Red and blue player can see each other if within a set distance; red attempts to deceive blue by blending in with white.

C

Red and blue player can see each other if within a set distance; red attempt to deceive blue by blending in with white; white informs on red with probability p if it sees red.

 

Mean

Standard Deviation

Experiment A

-28.62

9.0

Experiment B

-29.07

8.51

Experiment C, p=1

-28.09

8.1

Experiment C, p=0.5

-28.92

8.11

Learning curve for the blue player. Horizontal axis is the number of training epochs; Vertical axis is the mean of blue’s utility of 100 test episodes after each training epoch. Each epoch consists of 4000 game steps.

Varying Configurations Isolate the Effects of Deception

9

10 of 15

ANALYSIS

 

Red probabilistically sees blue

Blue probabilistically sees red

Red blends in with white

White acts as an informant

C1

probabilistic

C2

Within distance

probabilistic

C3

Within distance

Within distance

probabilistic

C4

Within distance

Within distance

probabilistic

C5

Within distance

Within distance

probabilistic

with probability p

 

Blue C1 model

Blue

C2 model

Blue

C3 model

Blue C4 model

Blue

C5 model

Red C1 model

-15.16

-14.94

-19.66

-24.84

-14.98

Red C2 model

-19.86

-14.52

-20.29

-25.45

-16.04

Red C3 model

-16.89

-11.75

-17.53

-22.75

-8.86

Red C4 model

-18.74

-14.20

-19.19

-24.30

-14.19

Red C5 model

-16.64

-11.93

-17.75

-23.08

-11.21

Varying Configurations Isolate the Effects of Deception

Game Scores for Varying Agent Pairs

Blue defender model trained to expect deception scores highest

10

11 of 15

EXAMPLES OF GAMES PLAYED (VIDEO)

11

12 of 15

CONCLUSIONS

  • Modeled and solved long-term competition between two opponents as a two-player plus a third-party player extensive form game
  • A capability to adapt courses of action to the strategy of an opponent able to use deception
  • Meta strategy-guided learning provides reward shaping, mitigates the sparse reward issue, and accelerates learning convergence
  • Deception (hiding) is embedded in the game – the optimal strategy uses deception when appropriate without the need for an additional mechanism
  • Readily extensible with other forms of deception, e.g., use of decoys

12

13 of 15

REFERENCES

  • [McMahan,Gordon,andBlum2003] McMahan, H. B.; Gordon, G. J.; and Blum, A. 2003. Planning in the presence of cost functions controlled by an adversary. In ICML’03.
  • [Fang, Stone, and Tambe 2015] Fang, F.; Stone, P.; and Tambe, M. 2015. When security games go green: Designing defender strategies to prevent poaching and illegal fishing. In IJCAI’15.
  • [Fang et al.2016] Fang,F.;Nguyen,T.H.;Pickles,R.;Lam,W.Y.; Clements, G. R.; An, B.; Singh, A.; Tambe, M.; and Lemieux, A. 2016. Deploying paws: Field optimization of the protection assistant for wildlife security. In AAAI’16.
  • [Xu et al. 2017] Xu, H.; Ford, B.; Fang, F.; Dilkina, B.; Plumptre, A.; Tambe, M.; Driciru, M.; Wanyama, F.; Rwetsiba, A.; and Nsubaga, M. 2017. Optimal patrol planning for green security games with black-box attackers. In GameSec’17.
  • [Wang et al. 2019] Wang, Y. ; Shi, Z. R.; Yu, L.; Wu, Y.; Singh, R.; Joppa, L.; and Fang, F. 2019. Deep Reinforcement Learning for Green Security Games with Real-Time Information. In AAAI’19.
  • [Lillicrap2016] Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; & Wierstra, D. 2016, Continuous Control With Deep Reinforcement Learning, ICLR 2016
  • [Lowe et al. 2017] Lowe, R.; Wu, W.; Tamar, A.; Harb, J.; Abbeel, P.; and Mordatch, I. 2017, Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments, https://arxiv.org/abs/1706.02275

13

14 of 15

Questions?

14

15 of 15

15