1 of 62

 Deep generative models for building virtual disease models and in-silico drug screening in complex diseases

Jun Ding, Ph.D.

Assistant Professor

FRQS Junior 2 Scholar

Endowed Meakins-Christe Chair in respiratory research

Meakins-Christie Labs, Department of Medicine

School of Computer Science

McGill University

MILA-Quebec AI Institute

07/22/2025

Google

2 of 62

3 of 62

Can we simulate diseases?

Quake et al, Cell, 2024

  1. Virtual cells
  2. Virtual disease (progression)
  3. Virtual drugs/perturbations

4 of 62

How to represent cells in-silico?

Lee, Jeongwoo, Do Young Hyeon, and Daehee Hwang. "Single-cell multiomics: technologies and data analysis methods." Experimental & Molecular Medicine 52.9 (2020): 1428-1442.

Sun, Yan V., and Yi-Juan Hu. "Integrative analysis of multi-omics data for discovery and functional studies of complex human diseases." Advances in genetics 93 (2016): 147-190.

Autoencoder

Autoencoder

PCA

5 of 62

I. Building virtual cells (cell representation learning)

Illustration generated with ChatGPT (OpenAI)

6 of 62

6

Gene

Gene Expression Value A

Image adapted from Andonix, 2020

Explore cells with the unexplored/under-utilized information

7 of 62

7

Gene

Gene Expression Value A

Image adapted from S. Kushal., 2023

Junction Reads

Exon Reads

Explore cells with the unexplored/under-utilized information

8 of 62

More Information

Step 1. Input Data Construction

Gene-level

Exon Graph

Cell-level

Exon Graph

Kailu Song

Conventional Gene Count Table

Kailu Song, Yumin Zheng, Bowen Zhao, David H. Eidelman, Jian Tang, Jun Ding. DOLPHIN advances single-cell transcriptomics beyond gene level by leveraging exon and junction reads, Nature Communications, 16 (6202) 2025

9 of 62

Step 2. Deep Generative Model

10 of 62

Step 3. Junction reads Aggregation

Boosting of Junction Reads

Cell Embeddings

Find Neighbors

(KNN)

11 of 62

Downstream Analysis

12 of 62

DOLPHIN Advances Cell Embedding Beyond Gene-Count Methods

Feature and adjacency matrices provide sufficient information for cell embedding.

DOLPHIN outperforms gene-count-based methods in cell embedding.

13 of 62

DOLPHIN Uncovers Pancreatic Cancer Markers Missed by Gene-Count Methods

14 of 62

MATES quantifies locus-specific TEs in single-cell data

Ignoring

Basic

Enhanced

Ruohan Wang#, Yumin Zheng#, Zijian Zhang#, Kailu Song, Erxi Wu, Xiaopeng Zhu, Tao P. Wu* & Jun Ding*. MATES: a deep learning-based model for locus-specific quantification of transposable elements in single cell. Nature Communications 15 (2024): 8798.

15 of 62

Multi-omic cell representation learning

Single-cell RNA-seq data:

× Can only reflect one side of a cell

✔ Great amount

Single-cell Multi-omics Data:

✔ Comprehensive view of a cell (More useful)

× Expensive and scarce

16 of 62

Xiuhui Yang, Koren K. Mann, Hao Wu* & Jun Ding*. scCross: a deep generative model for unifying single-cell multi-omics with seamless integration, cross-modal generation, and in silico exploration. Genome Biology 25.1 (2024): 1-34.

17 of 62

1. Train each modality independently with VAE

a. Utilize GCN to refine the feature and reduce the noise

b. Gene sets are inferred to supplement the biology information

18 of 62

2. Unify all trained VAEs with aligners to a common latent space

a. Modality integration via Probability distribution alignment

b. Modality integration via GAN

19 of 62

c. Modality alignment via MNN anchors

20 of 62

 

 

snmC-seq

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

scRNA-seq

scATAC-seq

Bi-directional architecture

Bi-direction architecture:

  1. Enable cross modal generation
  2. Achieve multi-omics simulation
  3. Realize multi-omics in-silico perturbation

21 of 62

Systematic benchmarks of integration performance

22 of 62

Effective integration of three omics layers

23 of 62

Single-cell cross modal generation

24 of 62

Cross modality generation

a, UMAP comparison (Intra-dataset) b, Cell type proportions comparison (Intra-dataset)

c, Expression levels comparison (Intra-dataset) d, Cell-cell interaction comparison (Intra-dataset)

25 of 62

Cross modality generation

26 of 62

Single-cell multi-omics simulation

27 of 62

Intra-modal simulation of matched single-cell multi-omics data

28 of 62

Large-scale single-cell sequencing remains expensive

Single-cell Gene Expression Data:

✔ High resolution Gene expression of each cell

× Expensive

Single-cell: $1.5m for 100 samples, �5000 cells per sample (costpercell)

Cells

Genes

Genes

Sample

Vs.

Single-cell (matrix/sample)

Bulk (vector/sample)

Bulk sequencing: $18k for 100 samples�(cost estimated at McGill Health Centre, 2023

29 of 62

Can we generate single-cell from more affordable bulk data?

Bulk Data:

✔ Affordable cost

× low resolution measurement (the average expression of all cells)

After

Enhancing the data

With deep generative AI

methods

Perfect Corp Homepage." Perfect Corp, https://www.perfectcorp.com/.

30 of 62

Single-cell semi-profiling pipeline

Jingtao Wang, Gregory J. Fonseca & Jun DingscSemiProfiler: Advancing large-scale single-cell studies through semi-profiling with deep generative models and active learning. Nature Communications 15.1 (2024): 5989.

31 of 62

Semi-profiled COVID-19 Cohort Displays Similar Visualizations to Real-profiled Cohort

32 of 62

Summary- Part I (cell representation)

1. Look into underexplored signals: Exploring underutilized signals in single-cell data enables more informative and robust cell representations.

2. Look beyond a single modality: Integrative multi-omic learning is critical for constructing comprehensive and biologically meaningful cell representations.

3. Look for cost-effective solutions: The high cost of single-cell sequencing remains a major bottleneck for large-scale cohort-level representation learning, underscoring the need for cost-efficient strategies.

33 of 62

II. How to represent disease progression and virtually screen for candidate therapeutics

Illustration generated with ChatGPT (OpenAI)

34 of 62

How to represent disease progression in-silico

Disease

  • Change of % of cellular population (from Healthy -> disease)
  • Change of Cellular state (e.g., transcriptome, epigenome, etc.)

Healthy

35 of 62

How to represent disease progression

……

36 of 62

IPF0 (Control)

IPF1 (Mild)

IPF2 (Moderate)

IPF3 (Severe)

Data: Single-cell time-series measurements of IPF lungs from different stages

How to represent disease progression (IPF is a progressive disease)

Robbie, Hasti, et al. "Evaluating disease severity in idiopathic pulmonary fibrosis." European Respiratory Review 26.145 (2017).

37 of 62

How to represent/model drugs in silico?

CMAP (https://clue.io/about)

Model drugs as perturbations of the model

38 of 62

Formalization: How to represent disease progression with AI/ML?�(i.e., in-silico model to represent disease progression)

IPF0 (Control)

IPF1 (Mild)

IPF2 (Moderate)

IPF3 (Severe)

Data: Single-cell time-series measurements of IPF lungs from different stages

Graphs�Nodes: cells/clusters

Edge: connections/transitions

39 of 62

How to to identify interventions that reverse or slow down the progression? �(perturb the in-silico drug disease progression model )

IPF0 (Control)

IPF1 (Mild)

IPF2 (Moderate)

IPF3 (Severe)

40 of 62

Why we need a new method? Is it necessary?

  • Existing deep learning-based methods
    • scGEN
    • chemCPA
    • GEARS

……

Limitations:

  • Need considerable labeled data in the training

(“Supervised”)

  • Many of these frameworks don’t particularly understand the disease

(“Generic”)

Lotfollahi, M., Wolf, F.A. and Theis, F.J., 2019. scGen predicts single-cell perturbation responses. Nature methods, 16(8), pp.715-721.

Hetzel, Leon, et al. "Predicting cellular responses to novel drug perturbations at a single-cell resolution." Advances in Neural Information Processing Systems 35 (2022)

Roohani, Y., Huang, K. and Leskovec, J., 2023. Predicting transcriptional outcomes of novel multigene perturbations with gears. Nature Biotechnology, pp.1-9.

41 of 62

 

VAE-GAN

to learn the cell embeddings

Graphical models (iDREM variant)

To learn the cell dynamics graph and

GRN

In-silico perturbation to score

and rank pathways/drugs

42 of 62

Single-cell data of the disease

  • 10 healthy donors (9 males and 1 female)
  • 9 IPF lung collected during transplantation surgeries (9 males)

  • Mesenchymal cell line
    • 231,544 cells

IPF snRNA-seq data from Naftali and Bart

43 of 62

Surrogate IPF temporal progression using different regions of IPF lung

  • Lung regions with different disease stages
    • 3 IPF affected stages
      • Normal looking (mild)
      • Moderate
      • Severe
  • Surface density (fibrosis quantification)
  • gaussian density estimator

severe

IPF lung

moderate

mild

mild

moderate

severe

healthy

McDonough, J. E. et al. Transcriptional regulatory model of fibrosis progression in the human lung. JCI Insight

4, e131597 (2019).

44 of 62

Employ a deep generative model to learn the cell embedding

45 of 62

How to build the dynamics graph?

Step 1: With the cell embeddings that we learn from the VAE-GAN model, we first cluster all the cells

Step 2: Starting from a terminal node (a cluster at the last stage), find its parent node in the previous stage and connect them with an edge => dynamics tree graph.

    • i. based on anchors (mutual nearest neighbors)
    • ii. based on KL divergence

In short, the components of the dynamics graph:

Nodes: cell clusters

Edges: connections between clusters between adjacent stages

46 of 62

How to infer the regulatory network underlying disease progression

Ernst, Jason, et al. "Reconstructing dynamic regulatory maps." Molecular systems biology 3.1 (2007): 74.

Ding, Jun, et al. "iDREM: Interactive visualization of dynamic regulatory networks." PLoS computational biology 14.3 (2018): e1006019.

Ding, Jun, et al. "Integrating multiomics longitudinal data to reconstruct networks underlying lung development." American Journal of Physiology-Lung Cellular and Molecular Physiology 317.5 (2019): L556-L568.

47 of 62

Emphasize the identified critical TFs and genes in the VAE-GAN model

48 of 62

In-silico perturbation strategies

  • Direct regulate target genes of drugs/compounds
  • Regulate genes according to the gene interaction networks

In-silico Perturbation

ii.

i

Gene A

Gene B

Gene C

Gene Interaction Networks

Direct Targets

Gene 1

Gene 2

Gene 3

Gene 4

Gene 5

Gene N

Raw

Expression

Perturbed

Expression

Drug

Impacts

49 of 62

In-silico perturbation scoring

  • Connectivity Maps (CMAP)
    • Impact of target genes of drugs

  • Perturb a cell population in a track
    • Quantify how much the perturbation can shift diseased cells towards healthier states
    • Perturbation score into [-1, 1]

 

Encoder

 

Latent Space

Perturbed

Stage 1

Control

Stage 1

Stage 2

Stage 3

Distance Original

Distance �Perturbed

 

 

 

 

 

 

 

 

Connectivity Maps

Perturbation Scores

Drug A

Drug B

Drug C

……

……

……

Subramanian, Aravind, et al. "A next generation connectivity map: L1000 platform and the first 1,000,000 profiles." Cell 171.6 (2017): 1437-1452.

50 of 62

Disease specific cell embeddings�

z

51 of 62

Cell embedding quality benchmarking

52 of 62

Ablation study- All components are essential

53 of 62

Iterative training empowers disease-focused cell embedding learning

iter_0

iter_1

iter_2

iter_3

iter_4

Silhouette score

0.12

0.14

0.16

0.18

0.20

54 of 62

UNAGI outperforms other benchmarked tools in simulated drug discovery

55 of 62

Critical pathways captured by UNAGI

Calcium signaling

Signaling by ROBO receptors

Collagens

TGF-beta signaling

ECM organization

Netrin-1 signaling

Syndecan 1 pathway

Lung fibrosis

Collagen formation

Signaling by GPCR

Perturbation Scores

FibAlv-4 Pathway

0.6

0.5

0.4

0.3

0.2

0.1

0.0

****

****

****

****

****

****

****

****

***

*

PMID: 33393489

PMID: 31393853

PMID: 25664495

PMID: 37005656

PMID: 32008852

PMID: 18161745

PMID: 18161745

PMID: 29293088

PMID: 34013369

56 of 62

UNAGI in-silico drug perturbations

  • Nintedanib (FDA approved)

  • Verify Nifedipine

PMID: 31367296

PMID: 24872318

PMID: 25656916

PMID: 25664495

PMID: 33671452

PMID: 35628257

57 of 62

Validate Nifedipine by Precise-cut lung slices (PCLS) experiments

  • PCLS experiments
    • DMSO (control)
    • Induce fibrosis (Fibrotic cocktail, FC)
    • Nifedipine and Nintedanib treatments

Day 5 cell embeddings

  • a control cocktail (CC) including all vehicles
  • a pro-fibrotic cocktail (FC) consisting of TGF-β (5 ng/ml, Bio-Techne),

PDGF-AB (10 ng/ml, Thermo Fisher), TNF-α (10 ng/ml, Bio-Techne), and LPA (5 µM, Cayman chemical)

58 of 62

PCLS experiments validate effectiveness of UNAGI unsupervised in-silico perturbation

  • Perform in-silico perturbation on Fibrotic cells

59 of 62

Significant alignment between real and in-silico drug perturbation

  • Real gene expressions: cells from real treatments
  • in-silico gene expressions: perturbed cells generated by UNAGI (in-silico treatment)

60 of 62

Benchmarking with methods in accuracy of in-silico perturbations

61 of 62

Summary-UNAGI

1) Unagi decodes the cellular dynamics from longitudinal single-cell data (disease in-silico)

2) Unagi empowers in-silico perturbations to identify candidate drugs (perturbation in-silico)

3) Unagi can be generalized to study other complex diseases such as COPD or CLAD

5) Unagi (for IPF) is freely available at https://github.com/mcgilldinglab/unagi

Yumin Zheng, Jonas C. Schupp, Taylor Adams, Geremy Clair, Aurelien Justet, Farida Ahangari, Xiting Yan, Paul Hansen, Marianne Carlon, Emanuela Cortesi, Marie Vermant, Robin Vos, Laurens J. De Sadeleer, Ivan O. Rosas, Ricardo Pineda, John Sembrat, Melanie Königshoff, John E. McDonough, Bart M. Vanaudenaerde, Wim A. Wuyts, Naftali Kaminski* & Jun Ding*A deep generative model for deciphering cellular dynamics and in silico drug discovery in complex diseases. Nature Biomedical Engineering (2025). https://www.nature.com/articles/s41551-025-01423-7

62 of 62

Thanks!

DingLab (McGill) students:

Yumin Zheng

Paul Hansen

Naftali Kaminski’s Lab at Yale

Bart Vanaudenaerde’s Lab at KU Leuven

Tao wu’s Lab at Baylor College of Medicine