Data splitting against information leakage
with DataSAIL
Roman Joeres, Drug Bioinformatics, HIPS/HZI
HelmholtzAI Conference, Karlsruhe, 03.06.2025
1
Information leakage in biomedical research
2
2
Information leakage in biomedical research
Bernett, Blumenthal, Grimm, Haselbeck, Joeres, Kalinina, List. Nature Methods (2024)
2
3
Drug-target interaction prediction models
Dataset
Model
Prediction
3
4
Common evaluation of drug-target interaction prediction models
random
cold-protein
cold-ligand
Random splits dominate but cold-splits are important
to test out-of-distribution performance
4
train
validation
test
5
Biochemical similarities enable different levels of coldness
similarity-based cold-protein
similarity-based cold-ligand
Biological molecules
are differently different.
Data splits have to account for that.
Myoglobin
Hemoglobin
Trypsin
5
6
What does a data splitter do?
Mathematically speaking:��Given dataset and sizes ��Find a partition of with
Solution:
There are many variants for :
Partitioning 1
Partitioning 2
Dataset
6
7
Existing algorithms fall short
Existing algorithms can have one or several shortcomings towards biological data:
one-dimensional data
two-dimensional data
custom data
7
8
Data splitting as a constraint optimization problem with DataSAIL
Objective:
are in
the same split
Minimize the inter-split similarities
8
similarity metric between data points
9
Data splitting as a constraint optimization problem with DataSAIL
Objective:
weights of a
data point
Minimize the inter-split similarities
8
10
Data splitting as a constraint optimization problem with DataSAIL
Objective:
subject to:
Minimize the inter-split similarities
Respect the requested split sizes
relative error for split sizes
8
11
Data splitting as a constraint optimization problem with DataSAIL
Objective:
subject to:
Minimize the inter-split similarities
Respect the requested split sizes
subset of D belonging to class c
relative error for class ratios
Keep the ratios between class
8
12
Data splitting as a constraint optimization problem with DataSAIL
Objective:
subject to:
Minimize the inter-split similarities
Respect the requested split sizes
This COP is NP-hard! The proof is in manuscript.
(Polynomial-time reduction from the Minimum k-Section Problem)
Keep the ratios between class
8
13
Full-scale DataSAIL pipeline
9
read any type of input ata
cluster the data
“solve” optimization problem
map results
14
Different types of splits for different types of datasets
10
15
Different types of splits for different types of datasets
10
16
Different types of splits for different types of datasets
10
17
Different types of splits for different types of datasets
10
18
Evaluation of DataSAIL
one-dimensional data
two-dimensional data
LP-PDBind:
Experimentally measured binding affinities of protein-ligand pairs
Tox21:
Predict the toxicity
of molecules in
different pathways
Datasets
Models
11
Random
Forest
SVM
Gradient
Boosting
MLP
DeepDTA
D-MPNN
for 2D data
for 1D data
19
Results for one-dimensional splits on Tox21
with stratification (here: active and inactive molecules in the SR-ARE assay)
without stratification
12
20
Results for two-dimensional splits on LP-PDBBind
This also works on the PLINDER benchmark,
better than the PLINDER algorithm
13
21
Summary & Outlook
David Blumenthal
FAU
Olga Kalinina
HIPS/UdS
Alex Gress, Anne Tolkmitt, Ilya Senatorov
all other colleagues and friends in the KalininaLab
and Daniel Bojar (Gothenburg University)
14
Acknowledgements
22
Summary & Outlook
David Blumenthal
FAU
Olga Kalinina
HIPS/UdS
Alex Gress, Anne Tolkmitt, Ilya Senatorov
all other colleagues and friends in the KalininaLab
and Daniel Bojar (Gothenburg University)
Objective:
subject to:
14
Acknowledgements
23
Summary & Outlook
David Blumenthal
FAU
Olga Kalinina
HIPS/UdS
Alex Gress, Anne Tolkmitt, Ilya Senatorov
all other colleagues and friends in the KalininaLab
and Daniel Bojar (Gothenburg University)
14
Acknowledgements
24
Summary & Outlook
Acknowledgements
David Blumenthal
FAU
Olga Kalinina
HIPS/UdS
Alex Gress, Anne Tolkmitt, Ilya Senatorov
all other colleagues and friends in the KalininaLab
and Daniel Bojar (Gothenburg University)
14
25
SI – Integer Linear Program
Minimize inter-split similarities
Respect the requested split sizes
Keep the ratios between class
All samples to exactly one split
Auxiliary bound for Eq. 1
Optimization variable
Auxiliary variable
15
26
SI – Runtime
16
27
SI – Limitations
17
28