1 of 8

Testing Consensus Protocols with Deep Reinforcement Learning

2 of 8

  1. https://medium.com/coinmonks/breaking-down-the-cosmos-game-of-stakes-5cbc538bcedb
  2. https://arxiv.org/abs/2004.10617
  3. https://arxiv.org/abs/1912.01798

Motivation

  • Formal Verification of consensus protocols is cumbersome
  • It is unclear how protocols behave when the byzantine agent is rational – what is the underlying objective?
  • A systematic way for the community to robustly test and find flaws in consensus protocols does not exist
    • Cosmos[1] : Game of Stakes
  • Related Work
    • Twins[2] : Brute force search running k instances of a node to simulate byzantine behavior
    • SquiRL[3] : Use DRL to recover known bitcoin attacks and find more optimal strategies in specific scenarios

3 of 8

Reinforcement Learning

  • Train agents to interact with an environment based off a given reward function
  • Agent explores state space in first few timesteps and then begins to exploit a strategy
  • Deep Reinforcement Learning – Using neural networks to model action probabilities and/or rewards

4 of 8

Simulation Environment

5 of 8

Simulation Environment

  • Multi-Agent Proximal Policy Optimization (PPO) Algorithm
  • Byzantine Agents share a 4-Layer MLP (256,128,128,64)
  • Honest Agents follow protocol specifications
  • Byzantine agents freely choose available actions
  • State: Vector of values agent receives from other agents
  • Run 100 epochs with 600-1000 simulations of the protocol in each epoch

6 of 8

Example Simulation

  • Rewards
    • Honest Correct Commit: -50
    • Delay Termination: 200
    • Violate Safety: 20

  • Rewards
    • Honest Correct Commit: -50
    • Delay Termination: 20
    • Violate Safety: 200

  • 3 Agents (1 Byzantine, 2 Honest)
  • Mutations: Remove Equivocation Checks

7 of 8

  1. https://eprint.iacr.org/2018/1028.pdf

Current State

  • Mutation Testing on Synchronous Byzantine Agreement[1]
    • Mutations: Removing a round of communication, removing equivocation checks from honest agents,
    • Reward byzantine agents based on delaying termination, violating safety, or both
    • Penalize byzantine agents if honest agents achieve consensus
  • Integrate action space with abstracted PKI
  • Train byzantine agents to attack both termination and safety

8 of 8

Key Issues and Future Plans

  • Remembering state and exploiting messages over longer time horizons
    • Updating policy/value networks to LSTM from MLP
  • Interact in environments under partial synchrony and asynchrony
  • Sparse Rewards – general DRL problem – to what extent should we give intermediate rewards to agents vs. give rewards at the end of a protocol rollout