1 of 20

Deep Learning Architectures for Automated Malaria Diagnostics

Dominik Polzer

2 of 20

Problem definition

3 of 20

4 of 20

5 of 20

Dominik

Polzer

6 of 20

Solution approach

Model ID

Size (MB)

Epochs

F1 Score

Technical Status

Model 0

4.04

10

98%

Baseline model (No augmentation)

Model 3

4.04

7

97%

Good Generalization (augmentation)

Model 4

58.20

3

~95%

Bouncy Validation (High Learning rate)

Model 2

18.95

9

~91%

Slight overfitting Issue

Model 5

58.29

5

~90%

Augmentation + PT

Model 1

18.92

5

FAIL

pre-processing fail

All models compared in the process of finding the best classification model.

7 of 20

Proposed model solution

97%

F1-SCORE PEAK

7 Epochs / 4.04MB

Model 3:

  • Lightweight & Memory-Efficient: ~1M parameters and 4.04 MB model size enable near-instant inference and deployment on edge, mobile, and offline medical devices.
  • Low Computational Overhead: Shallow CNN, efficient architecture requires less compute, making it ideal for real-time, battery-powered, and low-energy environments.
  • Robust in Real-World Settings: Trained with data augmentation, Model is resilient to image noise, rotation, and acquisition variability, supporting reliable clinical and field use.
  • Cost-Effective Deployment: No GPU or server-side infrastructure required, reducing deployment complexity and long-term operational costs.

8 of 20

Model 3 training process

9 of 20

Proposed business solution

10 of 20

11 of 20

12 of 20

Dominik Polzer

Appendix

More info on the hyperparameters chosen in the code -

shown in the next few slides …

13 of 20

Predictions - Confusion matrix

Model 3

14 of 20

Model predictions

We can observe some of the Model3 classification results.

15 of 20

False positives

16 of 20

Model 3 CNN Architecture

  • The model consists of 3 convolutional blocks, each combining Conv2D, MaxPooling, and Dropout for efficient feature extraction and regularization.

  • Feature maps are flattened and passed to a 2-layer dense classifier (512-unit hidden layer + 2-class softmax output).

  • In total, the architecture includes 13 layers, with 5 trainable (learnable) layers and ~1.06M parameters, optimized for lightweight deployment.

17 of 20

Pre-processing - Augmentation hyper-params

Model 3 training dataset preprocessing utilized these selected augmentation hyperparameters.

18 of 20

Comparison vs VGG16 architecture

  • Depth vs Efficiency
    • Unlike VGG16, which is a massive general-purpose image recognizer (138M params), Model 3 is a purpose-built specialist (~1M params). We traded depth for aggressive regularization (Dropout at every stage). This ensures the model remains lightweight enough for mobile phones (edge deployment) while avoiding overfitting, which is critical when dealing with medical imaging data that may have high variability.�
    • VGG16 original paper (Simonyan & Zisserman, 2014) only applied Dropout to the fully connected (dense) layers, not the convolutional blocks.�
  • Deployment Implications
    • VGG16 - 16 weight layers (13 Conv + 3 Dense) requires GPU acceleration and significant memory bandwidth.
    • Model 3 - 5 weight layers (3 Conv + 2 Dense) is optimized for CPU, edge, and mobile deployment, with far lower latency and power usage.

19 of 20

Top 4 observed models

Model training /fitting process.

Notice the sharp val_accuracy jumps in model 4 and 5.

This was likely caused by non optimal HIGH learning rate(0.001) hyper-parameter.

20 of 20

Resource references:

VGG 16 Architecture

WHO - Malaria report 2025

Dominik Polzer