EXPLORING CRYO-ET WITH DYNAMO: TILT SERIES ALIGNMENT, TEMPLATE MATCHING, AND PCA
RAFFAELE CORAY,
NUMERICAL METHODS OF CRYO ELECTRON TOMOGRAPHY LAB
MADRID, 2025
1
TILT SERIES ALIGNMENT IN DYNAMO
2
AUTOMATIC TILT SERIES ALIGNMENT
3
e-
e-
Idealization
Reality
50 nm
AUTOMATIC TILT SERIES ALIGNMENT
4
Micrograph i-2
Micrograph i-1
Micrograph i
Micrograph i+1
Micrograph i+2
chain
observations
AUTOMATIC TILT SERIES ALIGNMENT
5
Marker observation
Estimated
marker projection
Computation of marker projection
Indexing of observations
Model of the alignment geometry
AUTOMATIC TILT SERIES ALIGNMENT: DETECTION
6
AUTOMATIC TILT SERIES ALIGNMENT: INDEXING
7
Micrograph
Neighboring Micrograph
Initial matches create chains
AUTOMATIC TILT SERIES ALIGNMENT: INDEXING
8
Initial chains
Estimation in 3D
Reprojection
Re-Indexing
Incorrect indexing
Incomplete indexing
AUTOMATIC TILT SERIES ALIGNMENT: INDEXING
9
Missing chains
Apply current alignment
Reprojection
New chains
Reconstruction and inspection
10
HANDS-ON TUTORIAL
11
TEMPLATE MATCHING IN DYNAMO
12
PARTICLE PICKING
13
Template Matching
Manual
Automated
TEMPLATE MATCHING: OVERVIEW
14
Cross-correlation map
Assembled cross-correlation map
Best template rotations
Angle i
Angle i+1
Angle i+2
TEMPLATE MATCHING: HIGH RESOLUTION
15
200 nm
10 nm
TEMPLATE MATCHING: DIVIDE AND CONQUER
16
Template
Informations to reduced scanning requirements:
Spatial locality
Angular restrictions
TEMPLATE MATCHING: DIVIDE AND CONQUER
17
TEMPLATE MATCHING: IMPLEMENTATION
18
VIDEO?
On-the-fly reconstruction
Manual tesselation
Peak extraction
Chunk peak extraction
Single chunk modus
Automatic tesselation
Running time estimator
Model-aware template match module
HANDS-ON TUTORIAL
19
PRINCIPAL COMPONENT ANALYSIS IN DYNAMO
20
PCA: WHAT IS IT?
21
PCA is a dimensionality reduction technique.
… we cannot approximate it
with only one component
r
rx
ry
X
Y
Let us consider a data sets of points in the plane.
Each point has two components.
We can describe them with the ”regular”
x,y system of coordinates,
which is unaware of the data structure
if we consider a generic data point r
PCA: WHAT IS IT?
22
can be (better) approximated
by its projection along U, i.e. ru
U
V
ru
r
let’s try again.
Now we tailor a system of reference
for this data set
dimensionality reduction
now, our generic data point r
PCA: WHAT IS IT?
23
dimensionality
reduction
we go from 2 dimensions (x,y)
to 1 dimension (u)
How do we go from 256ˆ3 dimensions to, say, 10?
Each subtomogram has (for example) a sidelength of 256 voxels
In subtomogram averaging, each data point would be a particle
PCA: WHAT IS IT?
24
x2i
1005 | 1021 |
1021 | 1053 |
2029 | 0 |
0 | 29 |
Matrix diagonalization <-> redundancy reduction
PCA: CROSS-CORRELATION MATRIX
25
Missing wedge-aware cross-correlation coefficients for all possible particle couples
Element m(i,j)
similarity
of particle i
and particle j
Constrained
cross correlation
aligned particles
Förster et al, JSB, 2008
Castaño-Díez et al, JSB, 2012
aligned particles
PCA: MISSING WEDGE
26
each pair of particles are filtered to
their common Fourier components for a fair comparison!
“perfect” particle
Missing wedge compensation
when measuring particle similarity
PCA: MATRIX DIAGONALIZATION
27
A systematic way to create eigenvolumes as linear combinations of particles:
≈ p1,1 •
+ p1,2 •
≈ p2,1 •
≈ p3,1 •
≈ p4,1 •
+ p1,3 •
+ p2,2 •
+ p3,2 •
+ p4,2 •
+ p2,3 •
+ p3,3 •
+ p4,3 •
+ p1,N-1 •
+ p1,N •
+ p2,N-1 •
+ p3,N-1 •
+ p4,,N-1 •
+ • • •
+ • • •
+ • • •
+ • • •
+ p2,N •
+ p3,N •
+ p4,N •
L eigenvectors (eigenvolumes)
N particles
<<
PCA: A NEW BASIS TO REPRESENT THE PARTICLES
28
≈ λ1,1 •
+ λ1,2 •
+ λ1,3 •
+ λ1,4 •
≈ λN-2,1 •
≈ λ3,1 •
≈ λ2,1 •
≈ λN,1 •
≈ λ1N-1,1 •
+ λ2,2 •
+ λ3,2 •
+ λN-2,2 •
+ λN-1,2 •
+ λN,2 •
+ λ2,3 •
+ λ2,4 •
+ λ3,4 •
+ λN-2,4 •
+ λN-1,4 •
+ λN,4 •
+ λ3,3 •
+ λN-2,3 •
+ λN-1,3 •
+ λN,3 •
λi,j : “weight”
of eigenvolume j in particle I
“eigencomponents”
…
Particles
most relevant
eigenvector
less relevant
eigenvector
particles
eigenvectors
λi,j :
representation of the full data set as a single matrix
tailored to capture variance
across particles
PCA: A NEW REPRESENTATION TO IDENTIFY CLASSES
29
Each point represents a particle in a L-dimensional space
λ1,•
λ2,•
λ3,•
K-means: clustering on L-dimensional spaces provided by L Principal Components
Each axis represents
a column of the matrix
(i.e.: the distribution of a feature)
HANDS-ON TUTORIAL
30