1 of 26

Conversation Derailment Forecasting with Graph Convolutional Networks

YW

2 of 26

Graph Neural Network (GNN)

https://distill.pub/2021/gnn-intro/

3 of 26

Each layer is a GRAPH.

  • Image: MATRIX of pixels.
  • Language: SEQUENCE of vectors.
  • Graph:

Design Choice: V,E, (U) and their attributes.

https://distill.pub/2021/gnn-intro/

4 of 26

Graph connectivity doesn’t change.

Design Choice: f and activation function.

https://distill.pub/2021/gnn-intro/

5 of 26

Message Passing: Pooling

Design Choice: pooling function.

https://distill.pub/2021/gnn-intro/

6 of 26

Let’s put these together.

Design Choice:

  1. What to aggregate? Node, edge or both?

https://distill.pub/2021/gnn-intro/

7 of 26

Let’s put these together, cont’d

Design Choice:

Which should be aggregated first?

https://distill.pub/2021/gnn-intro/

8 of 26

More layers.

https://distill.pub/2021/gnn-intro/

Design Choice:

  1. How many layers?
  2. Final classifier?

9 of 26

 

AlexNet: https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf

10 of 26

Conversation Derailment

  • Off-topic/redirected conversation.

  • In this paper, derailment = personal attack.

11 of 26

Paper Profile

  • Research question: how to predict personal attack (PA) in a multi-party multi-turn conversation before the PA takes place?

  • Contribution:
    • More accurate prediction: in terms of a higher F1 score on 2 datasets.
    • Earlier prediction: 4 turns before the PA (forecast horizon)

  • Dataset:

12 of 26

Construction of CGN

  • Graph
    • Nodes
    • Connectivity
    • Edges
  • Network:
    • Message passing
    • Layers
    • Final classifier head

13 of 26

Graph-Nodes

  • t: textual input
    • BERT encoded

  • u: user id
    • Bi-LSTM encoded

  • s: score
    • number of up-vote – number of down-vote
    • Put into 6 bins: 3 for positive and 3 for negative
    • Bi-LSTM encoding.

14 of 26

Graph-Connectivity

15 of 26

Graph-Edges

  • Value: similarity-based attention module.

For each vertex, the incoming set of edges has a sum total weight of 1

16 of 26

Network

  • We use pooling to decide the value of each edge.

  • Pooling – Linear – ReLU – Linear - ReLU

17 of 26

Network: Classification Head

  • A graph-level task. All the vectors are used.

  • Fully-connected layer + Sigmoid activation

18 of 26

How do they train the model?

  •  

19 of 26

Finding: which is more important, t, s, u?

Larger Better.

20 of 26

How early can they detect a PA?

Larger Better.

21 of 26

My thoughts 1: Why do they need to change the graph?

  • Add more connectivity, otherwise pooling can’t be effective.

22 of 26

My thoughts 1: Why do they need to change the graph? Cont’d

  • However, there is another trick to do this, global node!

https://distill.pub/2021/gnn-intro/

23 of 26

My thoughts 2: How to decide the most important feature?

  • The classifier is a fully-connected layer + ReLU, so why not simply look at the weights of the linear layer?

24 of 26

My thoughts 3: Why static model has a better F score, but dynamic model has better forecast horizon?

  • Autor: “as it seems to be able to better model the dynamic relationships between the users of the turns with its graph model”
  • Me: this is all determined by the training data, no wonder.

25 of 26

My thoughts 4: Some surprising similarity

26 of 26

Thank you.

YW