1 of 16

Practical Data Harmonization in African Health Research

Presented by: Mr. Mugume Twinamatsiko Atwine

2 of 16

The Data Harmonization Challenge

3 of 16

Data Heterogeneity Within and Across DS-I Africa Research Hubs.

DS-I Africa's 13+ research hubs collect vast amounts of health data, but face critical challenges:

  • Inconsistent Variable Naming: Each DMAC uses different terminology for similar concepts (e.g., "BD018", "Country_01", "Country_02")
  • Missing Metadata: Many datasets lack proper variable descriptions, making interpretation difficult
  • Manual Mapping Bottleneck: Traditional harmonization requires tedious manual variable-by-variable mapping
  • FAIR Compliance Gap: Data often fails to meet Findable, Accessible, Interoperable, and Reusable standards
  • Time-Intensive Process: Months of work required to harmonize even small datasets

The result: Valuable health data remains siloed, limiting collaborative research potential across the consortium.

4 of 16

Typical Data Harmonization Process

5 of 16

Metadata Harmonization Workflow (Step 2 above)

6 of 16

The Data Harmonization Solutions

7 of 16

Metadata Harmonisation Tool: HE2AT.

8 of 16

Updated Metadata Harmonisation Tool: Local LLM Integration

Core Innovation:

  • Large Language Models (LLaMA 3.1:8b) for intelligent variable matching
  • Local Processing via Ollama - no data leaves your environment
  • Confidence Scoring for mapping recommendations with transparency

Core Organisation Issues:

  • Data Privacy: organisations are careful about data leakages.
  • Growing API budgets.
  • Rate Limits: no limitations on the amount of work you can do.
  • Customization control: you can adopt tools to different use cases.

9 of 16

How does the Tool Work?

10 of 16

Local LLM Platform:

11 of 16

Updated Streamlit Platform:

12 of 16

Updated Streamlit Platform:

13 of 16

Harmonization Process.

Content: Step 1: Upload Target Codebook - Define your harmonization standard.

Step 2: Upload Incoming Datasets - Study name, variables table, optional example data and protocols

Step 3: AI Description Generation - LLM extracts variable meanings from study documents using:

  • Vector embeddings for contextual understanding
  • Semantic retrieval of relevant documentation sections

Step 4: Recommendations - Algorithm suggests mappings using

  • Vector similarity calculations for semantic matching
  • DuckDB integration for fast recommendation retrieval

Step 5: Manual Validation - Human-in-the-loop quality control with AI assistance.

Step 6: Export Harmonized Mappings - Generate transformation files and harmonized datasets

14 of 16

Impact & Community Adoption

15 of 16

Accelerating Research Discovery Across DS-I Africa

Active Deployment:

  • Multi-Hub Testing: Successful pilots with MADIVA, HE2AT, and other research centers
  • Tool Accessibility: Available on GitHub with Docker deployment for easy setup

Next Integration Progress Objectives:

  • Data Transformation: data transformation integration in the current implementation
  • Data/Model Testing: Using different models to experiment what the best choices for these types of use cases.
  • Compiling Training Materials: Video tutorials, documentation, and creating best practices documents.

16 of 16

Thank You!