1 of 88

Tracking with Graph �Neural Networks

DANIEL MURNANE�BERKELEY LAB, CERN

1

Part 2: Extensions

HIGHRR LECTURE WEEK, HEIDELBERG UNIVERSITY

SEPTEMBER 13, 2023

HighRR Lecture Week - Heidelberg University - September 13, 2023

2 of 88

OVERVIEW

  • TrackML Competition & Dataset
  • TrackML Score & V-Score
  • Extending GNN4ITk: Faster, Better, Different
  • Faster…
    • Construction Upgrades
    • GNN Upgrades

2

  • Better…
    • Heterogeneity
    • Hierarchy
    • Checkpointing
  • Different…
    • Physics-motivated GNNs
    • Object condensation
    • … others?

HighRR Lecture Week - Heidelberg University - September 13, 2023

3 of 88

TRACKML COMPETITION & DATASET

3

HighRR Lecture Week - Heidelberg University - September 13, 2023

4 of 88

TRACKML CHALLENGE

  • A Kaggle Competition launched in 2018 for particle tracking with ML
  • “Generic detector” was used – ATLAS-like, but removed some of the complications: material effects, secondary particles, much of the noise, shared hits
  • Accuracy and throughput phases
  • Winners of each:

4

Arxiv:1904.06778

HighRR Lecture Week - Heidelberg University - September 13, 2023

5 of 88

TRACKML CHALLENGE

  • A Kaggle Competition launched in 2018 for particle tracking with ML
  • “Generic detector” was used – ATLAS-like, but removed some of the complications: material effects, secondary particles, much of the noise, shared hits
  • Accuracy and throughput phases
  • Winners of each:
    • Accuracy: TopQuarks – Uses seeds and track following. Conceptually similar to Kalman Filter
    • Throuput:

5

Arxiv:1904.06778

HighRR Lecture Week - Heidelberg University - September 13, 2023

6 of 88

TRACKML CHALLENGE

  • A Kaggle Competition launched in 2018 for particle tracking with ML
  • “Generic detector” was used – ATLAS-like, but removed some of the complications: material effects, secondary particles, much of the noise, shared hits
  • Accuracy and throughput phases
  • Winners of each:
    • Accuracy: TopQuarks – Uses seeds and track following. Conceptually similar to Kalman Filter
    • Throuput: Mikado – Also uses a similar concept to progress tracking, e.g. Kalman Filter
  • What’s the takeaway here? It’s not straightforward to beat the old ways!

6

Arxiv:2105.01160

HighRR Lecture Week - Heidelberg University - September 13, 2023

7 of 88

TRACKING METRICS

7

HighRR Lecture Week - Heidelberg University - September 13, 2023

8 of 88

TRACK MATCHING DEFINITIONS

  •  

8

 

Particle 1

Particle 2

Candidate 1

HighRR Lecture Week - Heidelberg University - September 13, 2023

9 of 88

TRACKML SCORE: WEIGHTED MATCHING

  1. Assign an importance to every hit in the event, which all sum to 1
    • Important hits: From long track, innermost and outermost hits, high pt hits
  2. A track is correctly “matched” to a particle if:
    • Strictly greater than 50% of hits in the track belong to that particle
    • Strictly greater than 50% of hits in the particle belong to that track

9

HighRR Lecture Week - Heidelberg University - September 13, 2023

10 of 88

TRACKML SCORE: WEIGHTED MATCHING

  1. Assign an importance to every hit in the event, which all sum to 1
    • Important hits: From long track, innermost and outermost hits, high pt hits
  2. A track is correctly “matched” to a particle if:
    • Strictly greater than 50% of hits in the track belong to that particle
    • Strictly greater than 50% of hits in the particle belong to that track

10

HighRR Lecture Week - Heidelberg University - September 13, 2023

11 of 88

TRACKML SCORE: WEIGHTED MATCHING

  1. Assign an importance to every hit in the event, which all sum to 1
    • Important hits: From long track, innermost and outermost hits, high pt hits
  2. A track is correctly “matched” to a particle if:
    • Strictly greater than 50% of hits in the track belong to that particle
    • Strictly greater than 50% of hits in the particle belong to that track

11

Particle 1

Particle 2

Candidate 1

Particle 1

Particle 2

Candidate 1

Particle 1

Particle 2

Candidate 1

Candidate 1

matched with

Particle 1

No match

No match

HighRR Lecture Week - Heidelberg University - September 13, 2023

12 of 88

TRACKML SCORE: WEIGHTED MATCHING

  1. Assign an importance to every hit in the event, which all sum to 1
    • Important hits: From long track, innermost and outermost hits, high pt hits
  2. A track is correctly “matched” to a particle if:
    • Strictly greater than 50% of hits in the track belong to that particle
    • Strictly greater than 50% of hits in the particle belong to that track
  3. All the weights of the matched hits are summed. A perfect matching of all tracks gives a sum of 1

12

HighRR Lecture Week - Heidelberg University - September 13, 2023

13 of 88

THE MATCHING PROBLEM

  • ATLAS-style? One-way? Two-way?
  • What percentage of hits matched?
  • Minimum number of hits in track and particle?
  • All particles equally important? All hits equally important?
  • What about shared hits? Can they be matched?

13

HighRR Lecture Week - Heidelberg University - September 13, 2023

14 of 88

EVALUATING A CLUSTERING

  • Assigning a single value to the “goodness of clustering” is non-trivial and non-obvious
  • Let’s consider an example to see the trade-offs
  • Consider a set of objects of type A and B

A

A

A

A

B

B

B

B

B

HighRR Lecture Week - Heidelberg University - September 13, 2023

15 of 88

EVALUATING A CLUSTERING

  • Assigning a single value to the “goodness of clustering” is non-trivial and non-obvious
  • Let’s consider an example to see the trade-offs
  • Consider a set of objects of type A and B
  • Let’s cluster them into cluster 1 and cluster 2

A

A

A

A

B

B

B

B

B

1

2

B

HighRR Lecture Week - Heidelberg University - September 13, 2023

16 of 88

EVALUATING A CLUSTERING

  • How should we measure our performance?
  • We can start by defining the entropy in each cluster

A

A

A

A

B

B

B

B

B

1

2

B

HighRR Lecture Week - Heidelberg University - September 13, 2023

17 of 88

EVALUATING A CLUSTERING

  •  

A

A

A

A

B

B

B

B

B

1

2

B

HighRR Lecture Week - Heidelberg University - September 13, 2023

18 of 88

EVALUATING A CLUSTERING

  •  

A

A

A

A

B

B

B

B

B

1

2

B

HighRR Lecture Week - Heidelberg University - September 13, 2023

19 of 88

EVALUATING A CLUSTERING

  •  

A

A

A

A

B

B

B

B

B

1

2

B

HighRR Lecture Week - Heidelberg University - September 13, 2023

20 of 88

HOMOGENEITY, COMPLETENESS AND V-SCORE

  • We can extend these ideas to capture the homogeneity and completeness across all clusters and all particles
  • The exact derivation is out-of-scope, but you should definitely look into mutual information to understand this properly!
  • At the end of the day:
    • Homogeneity is a measure of how well you’ve kept each cluster to a single particle type
    • Completeness is a measure of how well you’ve assigned all hits in a particle to a single cluster
  • These are the clustering analogy of purity and efficiency
  • However, Kaggle allows a single score to capture performance…

HighRR Lecture Week - Heidelberg University - September 13, 2023

21 of 88

HOMOGENEITY, COMPLETENESS AND V-SCORE

  •  

HighRR Lecture Week - Heidelberg University - September 13, 2023

22 of 88

EXTENSION: WEIGHTING THE V-SCORE

  • One final point: It is not equally important to cluster all points in a particle
  • If a particle leaves several high-energy hits, they should certainly be clustered together
  • If two particles have high energy hits, they should certainly not be clustered together
  • These leads us to create a new V-Score definition: the Weighted V-Score
  • The derivation is out-of-scope (the source code will be available at an upcoming version of scikit-learn)
  • Can deep dive if you’re interested

HighRR Lecture Week - Heidelberg University - September 13, 2023

23 of 88

EXTENSION: MANY-TO-MANY LABELLING

23

HighRR Lecture Week - Heidelberg University - September 13, 2023

24 of 88

THROUGHPUT & LATENCY

  • What is the goal?
  • Once moved to “offline tracking” essentially infinite time to reconstruct (although compute budget is limited)
  • In ATLAS HL-LHC trigger (aka “Event Filter”), have O(microseconds) time to reconstruct, maybe with some dip in efficiency
  • In some experiments (e.g. LHCb), aim to trigger on (essentially) all events, and perform on-the-fly full event reconstruction. In that case, target O(milliseconds) reconstruction with high accuracy

24

HighRR Lecture Week - Heidelberg University - September 13, 2023

25 of 88

SHORTCOMINGS OF GNN4ITK

25

HighRR Lecture Week - Heidelberg University - September 13, 2023

26 of 88

ACCURACY SHORTCOMINGS

  • Not perfect tracking efficiency
  • Poor performance in barrel strip modules

26

HighRR Lecture Week - Heidelberg University - September 13, 2023

27 of 88

THROUGHPUT SHORTCOMINGS

  • Physics is important, but GNNs shine in scaling behavior
  • When development began, graph-based pipeline started required 15 sec for TrackML
  • Implemented custom Fixed Radius Nearest Neighbor (FRNN) algo., cuGraph Connected Components algo., and Mixed Precision inference
  • Now have sub-second TrackML inference on 16Gb V100 GPU
  • Inference time scales approximately linearly across size of event, in TrackML

27

TrackML

HighRR Lecture Week - Heidelberg University - September 13, 2023

28 of 88

TRAINING COST SHORTCOMINGS

  • Even with the largest available GPUs (80Gb A100), still max out the memory with a relatively “small” GNN - 100k-1m parameters

28

  • What about if we want to go from spacepoints to clusters (300k nodes to 400k nodes), or to the next higher luminosity detector, or we want to train a very large GNN?

HighRR Lecture Week - Heidelberg University - September 13, 2023

29 of 88

FASTER GNN TRACKING

29

HighRR Lecture Week - Heidelberg University - September 13, 2023

30 of 88

FAST GRAPH CONSTRUCTION

30

  • Nearest neighbor search is a bottleneck of the graph construction stage
  • FAISS finding K=500 for N=100,000 ~ 700ms
  • KNN is overkill – we don’t need explicit list of K sorted neighbours
  • Built custom library on Fixed Radius Nearest Neighbour (FRNN) search algorithm
  • Cell-by-cell grid search is much faster: [The complexity of finding fixed-radius near neighbors. Bentley, et al 1977]

Fast fixed-radius nearest neighbors: Interactive Million-particle Fluids, Hoetzlein (NVIDIA), 2014

Accelerating NN Search on CUDA for Learning Point Clouds, Xue 2020

HighRR Lecture Week - Heidelberg University - September 13, 2023

31 of 88

FASTER SEGMENTATION

  • Many graph operations can be parallelised, and therefore are well-suited to GPU implementation
  • Connected components is one algorithm, which can be parallelised
  • Scipy has CPU version, which loops over each node with “Depth-first Search”
  • CuGraph searches many “frontiers” simultaneously

31

HighRR Lecture Week - Heidelberg University - September 13, 2023

32 of 88

FASTER HARDWARE

  • GPUs are used by default in our ML pipeline for training and inference
  • But FPGAs are a very low-latency option
  • Field Programmable Gate Array are able to compile a program to hardware, using Logic Elements and IO – essentially Look-up Tables (LUTs) that can capture any 4-input Boolean operators
  • Typically need to write functions from scratch, but HLS4ML is an effort to automatically compile Python ML frameworks to HLS (High Level Synthesis) language, then to the hardware language

32

HighRR Lecture Week - Heidelberg University - September 13, 2023

33 of 88

PRUNING

  • Hard to beat GPUs for big matrix multiplication – can be very efficiently multi-threaded
  • But large models typically only have a small subset of “important” weights (c.f. Lottery Ticket Hypothesis)
  • We can simply set those weights ~0 to exactly 0, but on GPU one still needs to run the full matrix multiplication
  • On FPGA, since the multiplication is in series, we can skip those 0 entries, and get a speed-up!

33

HighRR Lecture Week - Heidelberg University - September 13, 2023

34 of 88

QUANTIZATION

  • Similar to pruning, since everything is done manually on FPGA, we can choose how much precision we use to speed up
  • Can simply reduce precision of weights and operations after training
  • However there is significant improvement in performance using “Quantisation-aware Training” (QAT)

34

HighRR Lecture Week - Heidelberg University - September 13, 2023

35 of 88

QUANTIZATION

  • Similar to pruning, since everything is done manually on FPGA, we can choose how much precision we use to speed up
  • Can simply reduce precision of weights and operations after training
  • However there is significant improvement in performance using “Quantisation-aware Training” (QAT)

35

HighRR Lecture Week - Heidelberg University - September 13, 2023

36 of 88

MORE ACCURATE GNN TRACKING

36

HighRR Lecture Week - Heidelberg University - September 13, 2023

37 of 88

CHECKPOINTING

  • Graph construction leads to very large graphs O(1m) edges, cannot fit training on A100 GPU with 32Gb memory
  • Should not split the graphs up (leads to lower GNN accuracy)
  • Solution A: Were previously using a compromising form of “gradient checkpointing” – reduced memory by 4x
  • Now using maximal checkpointing, reduce memory further by 2x – just fits on A100

37

No checkpointing

Maximal checkpointing

Partial checkpointing

Graph

2

Graph Neural�Network

 

 

 

 

 

 

 

 

Edge Labeling

Edge Scores

HighRR Lecture Week - Heidelberg University - September 13, 2023

38 of 88

TRAINING SOLUTIONS

  • Solution B: Model offloading
  • Each layer of GNN placed on GPU for forward and backward pass, but held on CPU otherwise
  • Works well with TensorFlow, enabling training of O(1m) edge graphs
  • Unable to integrate with Pytorch pipeline

38

ZeRO-Offload: Democratizing Billion-Scale Model Training

arXiv: 2101.06840

Graph

2

Graph Neural�Network

 

 

 

 

 

 

 

 

Edge Labeling

Edge Scores

HighRR Lecture Week - Heidelberg University - September 13, 2023

39 of 88

BARREL STRIP MISCLASSIFICATION

39

ATLAS ITk

Nature of false positive edges

Location of false positive edges

43%: “True” ghosts

37%: Fakes

HighRR Lecture Week - Heidelberg University - September 13, 2023

40 of 88

BARREL STRIP MISCLASSIFICATION

40

Fake edges: 37%

Edges between SP from particle A and particle B. i.e. The GNN is “wrong”

“True” ghost edges: 43%

Edges between SP from particle A, and a ghost SP of clusters from particle A and particle B. I.e. The GNN is “right”, the construction is “wrong”

ATLAS ITk

 

 

Ghost

 

Location of false positive edges

HighRR Lecture Week - Heidelberg University - September 13, 2023

41 of 88

STRIP MODULES: GHOSTS AND�Z-RESOLUTION

  • Since spacepoints are constructed from pairs of clusters in the strip, could mis-construct and form a ghost
  • These ghosts can be cleaned up in later stages of the reconstruction chain
  • However, even for correctly matched clusters, there remains low z-resolution
  • Consider this example
  • Easily confuses GNN!
  • Could fix by including underlying cluster information somehow… (e.g. heterogeneous node features)

41

ATLAS ITk

Image courtesy of Jan Stark – thanks!

Cluster A

Cluster B

Constructed spacepoint

Ideal spacepoint

HighRR Lecture Week - Heidelberg University - September 13, 2023

42 of 88

CURRENT PIPELINE �PERFORMANCE

  •  

42

True Cluster A

True Cluster B

Constructed �spacepoint

Ideal spacepoint

Strip side A

Strip side B

HighRR Lecture Week - Heidelberg University - September 13, 2023

43 of 88

43

HighRR Lecture Week - Heidelberg University - September 13, 2023

44 of 88

IMPROVEMENT FROM INCLUDING CLUSTER INFORMATION

44

Only spacepoint information

Spacepoint+cluster information

HighRR Lecture Week - Heidelberg University - September 13, 2023

45 of 88

ONGOING WORK: HETEROGENEOUS NODE FEATURES

  •  

45

HighRR Lecture Week - Heidelberg University - September 13, 2023

46 of 88

ONGOING WORK: HETEROGENEOUS NODE FEATURES

  • To get intuition, consider simple filter �MLP applied to two pixel nodes:
  • To apply a filter MLP to a pixel (single cluster) and strip (double cluster) node combination, need a different MLP:

  • Already gives better than homogeneous filter MLP (~2x construction purity)

46

0

1

0

1

0

1

 

0

1

 

HighRR Lecture Week - Heidelberg University - September 13, 2023

47 of 88

ONGOING WORK: HETEROGENEOUS GRAPH NEURAL NETWORK

  •  

47

2

3

0

1

0

1

2

3

Node encoder 1

Edge encoder [1,1]

Edge encoder [0,1]

Node encoder 0

Edge encoder [0,0]

HighRR Lecture Week - Heidelberg University - September 13, 2023

48 of 88

HETEROGENEOUS GNN PERFORMANCE

  • The average total purity is 94% for both models
  • Adding model heterogeneity results in up to 11% improvement in GNN per-edge purity in the Strip barrel region, with ~1% loss in the Pixel subsystem

48

HighRR Lecture Week - Heidelberg University - September 13, 2023

49 of 88

HETEROGENEITY & THE MISSING HITS…

  • Now we can see why we are missing hits: there are orphan clusters in the strip that are never constructed into spacepoints
  • Would be great to have GNN that could handle both orphaned clusters and spacepoints…

49

HighRR Lecture Week - Heidelberg University - September 13, 2023

50 of 88

DIFFERENT APPROACHES TO GNN TRACKING

50

HighRR Lecture Week - Heidelberg University - September 13, 2023

51 of 88

HIERARCHY

51

HighRR Lecture Week - Heidelberg University - September 13, 2023

52 of 88

HIERARCHY

  • So far, our whole pipeline has been fairly “vanilla” (but it still took a lot of R&D to get this all to work!)
  • For example, every object in the graph is a spacepoint (or in the case of the heterogeneous GNN, either a strip spacepoint or pixel spacepoint)
  • But there are other granularities in the system, e.g. “track-like” objects

52

  • A hierarchical graph neural network is inspired by the different granularities of filter in a convolutional neural network

HighRR Lecture Week - Heidelberg University - September 13, 2023

53 of 88

HIERARCHY

  • Consider that in the GNN4ITk pipeline if a track is broken (a missing edge), there is no way to recover it

53

HighRR Lecture Week - Heidelberg University - September 13, 2023

54 of 88

HIERARCHY

  • Consider that in the GNN4ITk pipeline if a track is broken (a missing edge), there is no way to recover it
  • However, if we could “pool” hits together into track-like supernodes, then we could reconnect them at some other granularity

54

HighRR Lecture Week - Heidelberg University - September 13, 2023

55 of 88

HIERARCHY

  • Consider that in the GNN4ITk pipeline if a track is broken (a missing edge), there is no way to recover it
  • However, if we could “pool” hits together into track-like supernodes, then we could reconnect them at some other granularity
  • Can deep dive into this later if there’s interest. For now:

55

HighRR Lecture Week - Heidelberg University - September 13, 2023

56 of 88

SYMMETRY

56

HighRR Lecture Week - Heidelberg University - September 13, 2023

57 of 88

INCLUDING SYMMETRIES IN ML

  • The message passing in the GNN is “unconstrained” – any features can go in, and because of MLP Universal Approximator Theorem, any function may be learned that operates on input features
  • However, we can use our physics intuition to reduce the search space of this learned function
  • We know that there are symmetries in the geometry of some systems, which can be “built into” the GNN (and almost any other ML architecture)

57

HighRR Lecture Week - Heidelberg University - September 13, 2023

58 of 88

WHY EQUI-GNN: NAIVELY IMPROVING MODEL PERFORMANCE

58

Would love to add LundNet-5 and JEDI-net to this plot, but don’t have apples-to-apples rejection rate

HighRR Lecture Week - Heidelberg University - September 13, 2023

59 of 88

WHY EQUI-GNN: NAIVELY IMPROVING MODEL PERFORMANCE

59

Would love to add LundNet-5 and JEDI-net to this plot, but don’t have apples-to-apples rejection rate

NON-RELATIONAL ML

(SETS, IMAGES)

RELATIONAL ML

(GRAPHS)

PHYSICS-MOTIVATED ML

(SYMMETRY, DATA)

HighRR Lecture Week - Heidelberg University - September 13, 2023

60 of 88

WHY EQUI-GNN: NAIVELY IMPROVING MODEL PERFORMANCE

60

Would love to add LundNet-5 and JEDI-net to this plot, but don’t have apples-to-apples rejection rate

  • Given a particular ML structure (a.k.a relational bias), diminishing returns on simply increasing model size
  • Graph-structured appears to be as general as one can get structurally
  • GNN-based models seem to perform best at large size
  • Physics-based models seem to perform best at small size
  • Motivates us to constrain graph-structured ML with physics knowledge

HighRR Lecture Week - Heidelberg University - September 13, 2023

61 of 88

KINDS OF PHYSICS KNOWLEDGE

  • A variety of knowledge about the physics case can be included in the algorithm
  • Quantum field theory: Feynman diagram structure (EFP)
  • QCD: Decay processes in the Lund plane (LundNet)
  • Permutation invariance of the jet constituents (PFN, ParticleNet)
  • QCD + permutation invariance: Lund features with GNN (ParT: ParticleTransformer)
  • 2D translation invariance in the calorimeter (ResNeXt)
  • Special relativity: Frame-invariance under Lorentz transformations (LorentzNet, VecNet, Covariant ParT, …)

Good summary of theory-based tagging in Kasieczka, et al.

61

QFT Symmetries�

Spacetime �Symmetries�

Physics-informed �Features

Data�Augmentation

HighRR Lecture Week - Heidelberg University - September 13, 2023

62 of 88

WHAT DOES IT MEAN TO INCLUDE A SYMMETRY?

  •  

62

Track hits

 

Origin

 

 

HighRR Lecture Week - Heidelberg University - September 13, 2023

63 of 88

WHAT DOES IT MEAN TO INCLUDE A SYMMETRY?

  •  

63

Track hits

 

Origin

 

 

 

HighRR Lecture Week - Heidelberg University - September 13, 2023

64 of 88

WHAT DOES IT MEAN TO INCLUDE A SYMMETRY?

  •  

64

Track hits

 

Origin

 

 

 

HighRR Lecture Week - Heidelberg University - September 13, 2023

65 of 88

INVARIANCE VS. EQUIVARIANCE

  •  

65

HighRR Lecture Week - Heidelberg University - September 13, 2023

66 of 88

WHAT DOES IT MEAN TO INCLUDE A SYMMETRY?

  •  

66

0

2

1

3

4

 

 

 

 

0

2

1

3

4

 

 

 

 

HighRR Lecture Week - Heidelberg University - September 13, 2023

67 of 88

WHAT DOES IT MEAN TO INCLUDE A SYMMETRY?

  •  

67

0

2

1

3

4

 

 

 

 

0

2

1

3

4

 

 

 

 

HighRR Lecture Week - Heidelberg University - September 13, 2023

68 of 88

WHAT DOES IT MEAN TO INCLUDE A SYMMETRY?

  •  

68

0

2

1

3

4

 

 

 

 

0

2

1

3

4

 

 

 

 

HighRR Lecture Week - Heidelberg University - September 13, 2023

69 of 88

69

0

2

1

3

4

 

 

 

 

0

2

1

3

4

 

 

 

 

 

WHAT DOES IT MEAN TO INCLUDE A SYMMETRY?

Message passing invariant to �rotation and translation

Aggregation equivariant to rotation and translation

70 of 88

SO(2)-EQUIVARIANT GNN FOR TRACKING

  •  

70

HighRR Lecture Week - Heidelberg University - September 13, 2023

71 of 88

SO(2)-EQUIVARIANT GNN FOR TRACKING

  • This works, to a degree
  • Get good performance for very small models
  • At some point, an unconstrained model outperforms
  • Interestingly, even small unconstrained models learn the symmetry

71

HighRR Lecture Week - Heidelberg University - September 13, 2023

72 of 88

TRACKING AS OBJECT DETECTION

72

HighRR Lecture Week - Heidelberg University - September 13, 2023

73 of 88

THE TRACKING PROBLEM

  • Protons collide in center of detector, “shattering” into thousands of particles
  • The charged particles travel in curved tracks through detector’s magnetic field (Lorentz force)
  • A track is defined by the hits left as energy deposits in the detector material, when the particle interacts with material
  • In this study, we use the TrackML Dataset [link], with variable-sized subsets of tracks selected
  • The goal of track reconstruction: Given set of hits from particles in a detector, assign label(s) to each hit.

73

Can reframe the problem of assigning label 🡪 hits

  1. Assume the existence of some uniquely labelled “representative point” in each track object
  2. Then our task is to assign hits 🡪 representative point

1

1

1

1

2

3

4

5

6,7

8

9

2

3

3

3

4

4

4

5

5

5

6

7

6

7

6

7

8

8

8

9

9

9

1

2

3

4

5

6

8

9

7

Labels on hits

Hits to labels

HighRR Lecture Week - Heidelberg University - September 13, 2023

74 of 88

TRACKING AS OBJECT DETECTION

  • A well-studied problem in computer vision: Given an image, can we identify all discrete objects of interest and predict information about them?
  • Popular approach is to draw a bounding box as the representative label
  • Can’t directly use this approach for tracking: tracks are not localized in 3D space

74

The “You Only Look Once” (YOLO) approach to detection: draw a bounding box and predict the object in a single step.

Redmond et al, arXiv: 1506.02640

?

HighRR Lecture Week - Heidelberg University - September 13, 2023

75 of 88

OBJECT DETECTION AS METRIC LEARNING

  •  

75

Random hit 1

Random hit 2

Works quite well, but some points are clearly better candidates for representative than others. Can we learn which points are good representative points?

 

HighRR Lecture Week - Heidelberg University - September 13, 2023

76 of 88

OBJECT CONDENSATION: LEARNING REPRESENTATIVE POINTS

  •  

76

The potential function of members of the same class relative to the representation point of that class �(Kiesler 2020)

HighRR Lecture Week - Heidelberg University - September 13, 2023

77 of 88

DESIRED LOSS FUNCTION BEHAVIOUR: A TWITTER INSPIRATION

  • Idea: We can represent a social network as a directed graph of influence flow
  • Recuero et al, 2019, and Kim & Valente 2020 used network analysis to identify several types of user based on in-degree and out-degree of information flow
  • Let’s simplify: All members of network can be users (receive information from incoming edge) and influencers (send information to outgoing edge)
  • We can build a directed graph by learning for each member of the point cloud two embeddings in the same space: a user-embedding and an influencer-embedding

77

Kim & Valente 2020, COVID-19 Health Communication Networks on Twitter: Identifying Sources, Disseminators, and Brokers

Goal 1 �We would like users of each class to crowd around exactly one influencer that represents their class

Goal 2 �We want influencers to be distant from each other

HighRR Lecture Week - Heidelberg University - September 13, 2023

78 of 88

DESIRED LOSS FUNCTION BEHAVIOUR

  •  

78

 

Position of user-embeddings

Position of influencer-embeddings

 

 

In this case, 4 out of 5 users are in the neighbourhood

of an influencer

 

HighRR Lecture Week - Heidelberg University - September 13, 2023

79 of 88

DESIRED LOSS FUNCTION BEHAVIOUR

  •  

79

Position of user-embeddings

Position of influencer-embeddings

 

 

 

 

 

 

 

Case A

Case B

Case C

 

HighRR Lecture Week - Heidelberg University - September 13, 2023

80 of 88

THE INFLUENCER LOSS

  •  

80

 

 

The total Influencer Loss is at a minimum in this case

HighRR Lecture Week - Heidelberg University - September 13, 2023

81 of 88

A TRAINING MONTAGE

81

  • We can see the Influencer Loss working on two tracks above, across training epochs
  • In Real Space, we show only Users (circles) and Influencers (stars) when they are associated with an Influencer or User (respectively)
  • The color in Real Space is a projection in 1D of the location in Embedding Space
  • In Embedding Space, we should edges created, and connected Influencers are large stars, unconnected Influencers are small stars

HighRR Lecture Week - Heidelberg University - September 13, 2023

82 of 88

A TRAINING MONTAGE

82

REAL SPACE

EMBEDDING SPACE

  • We can see the Influencer Loss working on two tracks above, across training epochs
  • In Real Space, we show only Users (circles) and Influencers (stars) when they are associated with an Influencer or User (respectively)
  • The color in Real Space is a projection in 1D of the location in Embedding Space
  • In Embedding Space, we should edges created, and connected Influencers are large stars, unconnected Influencers are small stars

HighRR Lecture Week - Heidelberg University - September 13, 2023

83 of 88

GNN TRACKING IN PRODUCTION

83

HighRR Lecture Week - Heidelberg University - September 13, 2023

84 of 88

CONVERSION TO ONNX

  • Onnx is the gold standard for portability of ML
  • Takes any framework (Tensorflow, Pytorch, Jax)
  • Represents as a computational graph
  • However, graph neural network operations have always been lacking in Onnx
  • The latest version of Pytorch operations and Onnx libraries supports GNN conversion!

84

HighRR Lecture Week - Heidelberg University - September 13, 2023

85 of 88

CONVERSION TO C++

  • This can be done with Onnx, with C++ library of OnnxRuntime
  • Works basically out-of-the-box!
  • Can also do this with LibTorch

85

HighRR Lecture Week - Heidelberg University - September 13, 2023

86 of 88

OPEN PROBLEMS

86

HighRR Lecture Week - Heidelberg University - September 13, 2023

87 of 88

OPEN PROBLEMS

  •  

87

TrackML

ATLAS ITk

HighRR Lecture Week - Heidelberg University - September 13, 2023

88 of 88

OPEN PROBLEMS

  • We have typically been afraid of “dense” representations, hence the building of more and more sparse graphs. But sparse representations may miss interesting relationships, and models like GOAT try to apply dense models to graphs
  • We cannot yet get GNNs onto FPGAs easily – indexing and scattering is non-trivial
  • Training GNNs across GPUs is non-trivial:
    • Event-level parallelism gives no benefit (typical LHC hit-graph above OpenAI model of noise scale)
    • Node-level parallelism is difficult to implement (but we are working with Georgia Tech group to do this, library called Lasagne has been successfully tested with parallelised graph attention network)
  • Currently use spacepoints as the lowest level object, but ATLAS tracking natively uses clusters – should enable this in GNN pipeline

88

HighRR Lecture Week - Heidelberg University - September 13, 2023