1 of 30

EXPLORING CRYO-ET WITH DYNAMO: TILT SERIES ALIGNMENT, TEMPLATE MATCHING, AND PCA

RAFFAELE CORAY,

NUMERICAL METHODS OF CRYO ELECTRON TOMOGRAPHY LAB

MADRID, 2025

1

2 of 30

TILT SERIES ALIGNMENT IN DYNAMO

2

3 of 30

AUTOMATIC TILT SERIES ALIGNMENT

3

e-

e-

Idealization

Reality

50 nm

4 of 30

AUTOMATIC TILT SERIES ALIGNMENT

4

Micrograph i-2

Micrograph i-1

Micrograph i

Micrograph i+1

Micrograph i+2

chain

observations

5 of 30

AUTOMATIC TILT SERIES ALIGNMENT

5

Marker observation

Estimated

marker projection

Computation of marker projection

Indexing of observations

Model of the alignment geometry

6 of 30

AUTOMATIC TILT SERIES ALIGNMENT: DETECTION

6

7 of 30

AUTOMATIC TILT SERIES ALIGNMENT: INDEXING

7

Micrograph

Neighboring Micrograph

Initial matches create chains

8 of 30

AUTOMATIC TILT SERIES ALIGNMENT: INDEXING

8

Initial chains

Estimation in 3D

Reprojection

Re-Indexing

Incorrect indexing

Incomplete indexing

9 of 30

AUTOMATIC TILT SERIES ALIGNMENT: INDEXING

9

Missing chains

Apply current alignment

Reprojection

New chains

Reconstruction and inspection

10 of 30

10

11 of 30

HANDS-ON TUTORIAL

11

  • Walkthrough on manual marker clicking
    • Familiarization to observation/chain logic.
    • Tool useful for manual corrections.

  • Walkthrough on GUI based tilt series alignment
    • Familiarization to the automatic alignment algorithm.
    • Results visualization by GUI and figures.

  • Walkthrough on command line based tilt series alignment
    • Example of scripting for large dataset processing.

12 of 30

TEMPLATE MATCHING IN DYNAMO

12

13 of 30

PARTICLE PICKING

13

Template Matching

Manual

Automated

14 of 30

TEMPLATE MATCHING: OVERVIEW

14

Cross-correlation map

Assembled cross-correlation map

Best template rotations

Angle i

Angle i+1

Angle i+2

15 of 30

TEMPLATE MATCHING: HIGH RESOLUTION

15

200 nm

10 nm

16 of 30

TEMPLATE MATCHING: DIVIDE AND CONQUER

16

Template

Informations to reduced scanning requirements:

Spatial locality

  • chunks

Angular restrictions

  • cones

17 of 30

TEMPLATE MATCHING: DIVIDE AND CONQUER

17

18 of 30

TEMPLATE MATCHING: IMPLEMENTATION

18

VIDEO?

On-the-fly reconstruction

Manual tesselation

Peak extraction

Chunk peak extraction

Single chunk modus

Automatic tesselation

Running time estimator

Model-aware template match module

19 of 30

HANDS-ON TUTORIAL

19

  • Walkthrough for template matching
    • Familiarization to syntax and parameters.
    • Creation of template.
    • Parallel computing and chunks.
    • Analysis of cross-correlation maps.
    • Selection of TM particles.

  • Walkthrough for GUI-based template matching
    • Familiarization with GUI tools for tessellation creation.
    • Tomogram investigation with template matching.
    • Use of tessellation for template matching speed-up.

20 of 30

PRINCIPAL COMPONENT ANALYSIS IN DYNAMO

20

21 of 30

PCA: WHAT IS IT?

21

PCA is a dimensionality reduction technique.

… we cannot approximate it

with only one component

r

rx

ry

X

Y

Let us consider a data sets of points in the plane.

Each point has two components.

We can describe them with the ”regular”

x,y system of coordinates,

which is unaware of the data structure

if we consider a generic data point r

22 of 30

PCA: WHAT IS IT?

22

can be (better) approximated

by its projection along U, i.e. ru

U

V

ru

r

let’s try again.

Now we tailor a system of reference

for this data set

dimensionality reduction

  • an axis U along the direction of maximal variance
  • an axis V orthogonal to U

now, our generic data point r

23 of 30

PCA: WHAT IS IT?

23

dimensionality

reduction

we go from 2 dimensions (x,y)

to 1 dimension (u)

How do we go from 256ˆ3 dimensions to, say, 10?

Each subtomogram has (for example) a sidelength of 256 voxels

In subtomogram averaging, each data point would be a particle

  • pay a price: approximation

  • figure out a data-aware representation

24 of 30

PCA: WHAT IS IT?

24

x2i

 

 

1005

1021

1021

1053

 

 

2029

0

0

29

Matrix diagonalization <-> redundancy reduction

25 of 30

PCA: CROSS-CORRELATION MATRIX

25

Missing wedge-aware cross-correlation coefficients for all possible particle couples

Element m(i,j)

similarity

of particle i

and particle j

Constrained

cross correlation

aligned particles

Förster et al, JSB, 2008

Castaño-Díez et al, JSB, 2012

aligned particles

26 of 30

PCA: MISSING WEDGE

26

each pair of particles are filtered to

their common Fourier components for a fair comparison!

“perfect” particle

Missing wedge compensation

when measuring particle similarity

27 of 30

PCA: MATRIX DIAGONALIZATION

27

A systematic way to create eigenvolumes as linear combinations of particles:

≈ p1,1

+ p1,2

≈ p2,1

≈ p3,1

≈ p4,1

+ p1,3

+ p2,2

+ p3,2

+ p4,2

+ p2,3

+ p3,3

+ p4,3

+ p1,N-1

+ p1,N

+ p2,N-1

+ p3,N-1

+ p4,,N-1

+ • • •

+ • • •

+ • • •

+ • • •

+ p2,N

+ p3,N

+ p4,N

L eigenvectors (eigenvolumes)

N particles

<<

28 of 30

PCA: A NEW BASIS TO REPRESENT THE PARTICLES

28

≈ λ1,1

+ λ1,2

+ λ1,3

+ λ1,4

≈ λN-2,1

≈ λ3,1

≈ λ2,1

≈ λN,1

≈ λ1N-1,1

+ λ2,2

+ λ3,2

+ λN-2,2

+ λN-1,2

+ λN,2

+ λ2,3

+ λ2,4

+ λ3,4

+ λN-2,4

+ λN-1,4

+ λN,4

+ λ3,3

+ λN-2,3

+ λN-1,3

+ λN,3

λi,j : “weight”

of eigenvolume j in particle I

“eigencomponents”

Particles

most relevant

eigenvector

less relevant

eigenvector

particles

eigenvectors

λi,j :

representation of the full data set as a single matrix

tailored to capture variance

across particles

29 of 30

PCA: A NEW REPRESENTATION TO IDENTIFY CLASSES

29

Each point represents a particle in a L-dimensional space

λ1,

λ2,

λ3,

K-means: clustering on L-dimensional spaces provided by L Principal Components

Each axis represents

a column of the matrix

(i.e.: the distribution of a feature)

30 of 30

HANDS-ON TUTORIAL

30

  • Walkthrough on PCA through the command line
    • PCA computation: particle alignment, constrained cross-correlation matrix, matrix diagonalization.
    • Management of computing resources.
    • Inspection of eigencomponent and eigenvolumes.
    • Munual and automatic clustering.
    • GUI-based inspection of classes