1 of 41

Dimagi’s contribution to Equitable AI

Leveraging Large Language Models �to Address Global Inequalities

2 of 41

Large Language Models (LLMs) exhibit language abilities previously seen only from humans:

  • Adaptability, common sense, creativity, empathy, intuition, problem solving.

Likely to be transformative in coming years:

  • Likely faster than the 1-2 decades that recent digital advances have taken.

Tremendous risk and opportunity:

  • Our collective challenge is to make adoption of LLMs impactful, equitable, safe, and inclusive.

Perspective on Large Language Models

3 of 41

Discovery of potential impactful use cases: While many are focused on Q&A bots, Dimagi is exploring Chatbot-led interactions such as coaching, role play, interviewing, and more.

Platform: We are building Open Chat Studio to make it possible rapidly build and deploy chatbots in LMICs. This accelerates our work and is an open source resource for others.

Community: We are sharing our platform and learnings with peers and partners. Dimagi hosted a webinar in February with 240+ signups, and currently has 45+ active orgs on the Open Chat Studio platform.

Evidence base: Our next steps are to collaboratively build an evidence base on the impact of LLMs for global development.

Focus on FLWs: Dimagi experience and technology makes us especially well poised to leverage LLM to support frontline work.

Dimagi’s Contribution to Equitable AI

4 of 41

Dimagi’s exploratory, field-shaping AI portfolio

Individuals

How might we design AI to be used by anyone, anywhere?

Direct to Consumer

Chatbots

Frontline workers

How might AI support better jobs for better outcomes?

Personal AI Coaches

Ecosystem

How might we create an inclusive AI developer ecosystem?

For funding and

implementing partners

5 of 41

Dimagi’s Portfolio of LLM Projects

Open Chat Studio Projects (>$6m total funding)

Country

Partner

Funder

For Individuals

Family Planning Agency

Kenya, Senegal

Shujaaz, RAES

BMGF

Substance Use Disorder support

US

MGH, UCLA

NIH

Adolescent Mental Health Support

US

Columbia University

Columbia University

TB Education

Netherlands

KNCV TB

SMT

For Frontline Workers

Support for Group Chat

Somalia

Save the Children

Save the Children

Serious Illness Conservations

US

Ariadne Labs

NIH

CHW Coaching in Chichewa

Malawi

PACHI

BMGF GC

For Ecosystem Partners

Concept Note Review

General

BMGF

BMGF

GenAI Assistants for program managers

Nigeria

WHO / ESPEN

BMGF

Diabetes Health Coach

Global

Sanofi

Sanofi

6 of 41

Open Chat Studio

7 of 41

Platform for Building LLM-based chatbots

Open Chat Studio*

Large Language Model

(e.g., GPT-4)

*Initial development supported by BMGF FP investment on conversational agents in Senegal and Kenya

User interface

(e.g., Web, Telegram, WhatsApp)

8 of 41

Open Chat Studio: Harnessing the Power of LLMs

  • Make Your Own Chatbots: With Open Chat Studio, you can easily create your own chatbots using Large Language Models. OCS is built for use by program staff, M&E and other teams - you don't need to be an engineer.

  • Launch and Share: Use Open Chat Studio to deploy your chatbots on the web and via mobile apps such as Telegram and WhatsApp, with more options coming soon. Chatbot users do not need to have an account on Open Chat Studio to use your chatbots.

  • Access and Analyze Data: View and export detailed interaction data from your chatbots, conveniently formatted in CSV for further analysis.

Dimagi is now able to offer organizations free trial use of Open Chat Studio to create and deploy their own bots. Please reach out to ocs-info@dimagi.com to create a team space for your organization.

9 of 41

Open Chat Studio: Current Features

End user interface

  • WhatsApp, Telegram
  • Transcribe voice notes through OpenAI Whisper
  • Text to speech for LLM output using Microsoft Azure [Example Kangaroo Mother Care Bot interaction ]

Design guardrails

  • Safety layers
  • User input modification
  • Source material

Bot features for end users

  • Timers / proactive messages
  • Longitudinal bots

Research

  • Consent forms
  • Pre/post surveys
  • Review/Summary of transcripts

Advanced features

  • Leveraging OpenAI’s Assistants API for advanced knowledge retrieval and other functions.

10 of 41

Sample Chatbot Prompt

You are a personal coach talking to women in Tanzania who are trying to return to the workforce after having a child. Speak in a warm and nurturing tone. Speak in short sentences and limit each message to 2-3 lines. Introduce yourself and your purpose. Ask the user how their day was and give an empathetic response.

Ask the user what they'd like help with today. Provide examples: e.g., ask if they want coaching on how to get a new job, to learn a new resilience skill, or how to talk to their partner about going back to work. Based on how they respond:

  • You can offer to do a role-play to help them practice talking to their partner about returning to work. In the role-play, pretend to be their partner who is ambivalent or discouraging. Give the user feedback after three exchanges.
  • You can teach a short resilience practice to help address ongoing worries that may inhibit them, for example worries about being unsuccessful and the financial implications of that on their family.
  • You can offer to help them practice for a job interview for a job of their choice.

11 of 41

Prompt Builder

12 of 41

3

Why we’ve focused on GPT-4 so far

OpenAI’s GPT-4 much more powerful than other LLM we have tested.

Dimagi is using GPT-4 as a proxy for what open source solutions will be able to do relatively soon.

Using open source LLMs is likely the clearest path to privacy, data sovereignty, cost effectiveness, etc.

Open Chat Studio is as fast, secure, low-cost, and reliable as the LLM it uses.

13 of 41

Examples Chatbots

14 of 41

Chatbot Purpose: To discuss use of Multiple Micronutrient Supplements (MMS) with pregnant women to understand use and barriers they face.

LLMs: allow flexible, intuitive interaction.

Open Chat Studio Functionality:

  • Recording transcripts
  • Deploying via multiple channels: Web, Telegram, WA, etc.
  • Input formatting to adhere to purpose
  • Safety guardrails
  • Use of voice

[Example voice interaction with MMS Survey bot: ]

Client-facing Bot to Ask About MMS Usage

For more information on this bot, click the links below:

  • Read the MMS-survey chatbot prompt
  • Try the MMS-survey chatbot demo!
  • Read the full chatbot transcript

15 of 41

Chatbot Purpose: To discuss use of Multiple Micronutrient Supplements (MMS) with pregnant women to understand use and barriers they face.

LLMs: allow flexible, intuitive interaction.

Open Chat Studio Functionality:

  • Recording transcripts
  • Deploying via multiple channels: Web, Telegram, WA, etc.
  • Input formatting to adhere to purpose
  • Safety guardrails
  • Use of voice

[Example voice interaction with MMS Survey bot: ]

Client-facing Bot to Ask About MMS Usage

For more information on this bot, click the links below:

  • Read the MMS-survey chatbot prompt
  • Try the MMS-survey chatbot demo!
  • Read the full chatbot transcript

16 of 41

Can change output by adding the following to the bot’s instructions:

Speak in a way suitable for SMS. Make sentences as short as possible by using text speak (e.g., great becomes gr8). Use relevant emojis.

Client-facing Bot to Ask About MMS Usage

For more information on this bot, click the links below:

  • Read the MMS-survey for SMS chatbot prompt
  • Try the MMS-survey for SMS chatbot demo!
  • Read to full SMS chatbot transcript

17 of 41

Can change output by adding the following to the bot’s instructions:

Speak in Swahili.

Client-facing Bot to Ask About MMS Usage

For more information on this bot, click the links below:

18 of 41

Can change output by adding the following to the bot’s instructions:

Speak in Sheng. Learn to construct your questions and responses by using example Sheng conversations supplied here [~600 words of example Sheng conversation followed]

Client-facing Bot to Ask About MMS Usage

For more information on this bot, click the links below:

19 of 41

Chatbot Purpose: Give provider practice talking to women reluctant to take MMS.

Many other use-cases for role-play, e.g.:

  • Negotiating with husband or mother-in-law about birth spacing, work outside the home
  • Practice talking to provider to request contraceptives

Need for evidence: this is a great example of a potentially impactful use case that requires careful evaluation to assess usefulness.

FLW Coach for Supporting MMS

For more information on this bot, click the links below:

  • Read to FLW Coach for MMS chatbot prompt
  • Try the FLW Coach for MMS chatbot demo!
  • Read the full FLW Coach for MMS chatbot transcript

20 of 41

Chatbot Purpose: Give provider practice talking to women reluctant to take MMS.

Many other use-cases for role-play, e.g.:

  • Negotiating with husband or mother-in-law about birth spacing, work outside the home
  • Practice talking to provider to request contraceptives

Need for evidence: this is a great example of a potentially impactful use case that requires careful evaluation to assess usefulness.

FLW Coach for Supporting MMS

For more information on this bot, click the links below:

  • Read to FLW Coach for MMS chatbot prompt
  • Try the FLW Coach for MMS chatbot demo!
  • Read the full FLW Coach for MMS chatbot transcript

21 of 41

Chatbot feedback from following prompt:

The moderator will offer 2 examples of positive feedback on effective elements of the conversation, giving examples; and, suggest 2 areas for improvement, also providing examples from the conversation.

FLW Coach for supporting MMS - Feedback

For more information on this bot, click the links below:

  • Read to FLW Coach for MMS chatbot prompt
  • Try the FLW Coach for MMS chatbot demo!
  • Read the full FLW Coach for MMS chatbot transcript

22 of 41

Chatbot purpose: Summarize and discuss findings from interactions with clients and FLWS

Open Chat Studio Functionality:

  • Load and upload files
  • [On roadmap] Combining multiple bots to work together

Data Analysis for Program Improvement

For more information on this bot, click the links below:

  • Read the MMS Summary Bot chatbot prompt
  • Try the MMS Summary Bot chatbot demo!
  • Read the full MMS Summary Bot chatbot transcript (lightly edited on slide)
  • Note: Transcripts used to in order to generate this summary bot demonstration were fabricated by GPT (fake survey data was summarized.)

23 of 41

Chatbot Purpose: Daily check-in with pregnant client to understand and encourage MMS use.

Open Chat Studio Functionality:

  • Prompting user after time passes
  • Time awareness
  • Summarization of past context
  • [On roadmap] Ability to not respond to user

Tech platform example: Longitudinal bots

For more information on this bot, click the links below:

  • Read the Daily MMS Tracker Bot chatbot prompt
  • Try the to Daily MMS Tracker Bot chatbot demo!
  • Note: The transcript on this slide is an illustrative example, not actual output from bot linked above. Not all functionality included in this illustrative example is currently available on OCS.

24 of 41

Results:

Low Resource Languages

25 of 41

Baseline Chatbot: A role-play bot that helps Community Health Volunteers (CHVs) practice challenging conversations with vaccine hesitant clients.

Variations: We tested five different combinations of the following techniques for getting the chatbot to speak in Chichewa:

  • Provide prompt in Chichewa as well as English
  • Tell bot to speak in simple language, at primary school level
  • Lower GPT temperature parameter to 0.5 rather than 0.7
  • Tell bot to speak in local dialect Chewa

Transcript Generation: On Nov 22, 2023, 22 CHVs were given one of the bots to test. This yield 18 usable transcripts.

Transcript Rating: On Dec 13, 2023, 22 different CHVs (uninvolved in transcript generation) were each asked to rate the quality of the Chichewa on 4 transcripts using a 13-question survey. During data cleaning, six transcripts were rejected, leaving 82 for analysis.

Language Evaluation Methodology

Implemented with Pachi in Malawi and with funding from Grand Challenges for Catalyzing Equitable AI Use

26 of 41

Malawian Chichewa Language and Proficiency

Technique

Num ratings

Please rate the Chichewa language of the bot and its ability to communicate in Chichewa on a scale of 1-5.

1- Very Poor

2- Poor

3- Acceptable

4- Good

5- Very Good

When looking at the bot’s responses, it seems like a person who is:

1- Just learning how to speak Chichewa

2- Somewhere further along the journey of speaking Chichewa

3- Good at conversing in Chichewa

4- Excellent, a native Chichewa speaker

Prompt in English

17

4.00 [3.43, 4.57]

2.88 [2.41, 3.36]

Prompt in English and Chichewa

17

3.94 [3.52, 4.37]

2.71 [2.20, 3.21]

Prompt in English and Chichewa

Simple language

16

4.25 [3.89, 4.61]

3.19 [2.84, 3.54]

Prompt in English and Chichewa

Simple language

Temperature reduced to 0.5

16

4.38 [3.90, 4.85]

3.06 [2.65, 3.47]

Prompt in English

Simple language

Speak in local dialect (Chewa)

16

4.06 [3.57, 4.56]

3.19 [2.70, 3.67]

27 of 41

Malawian Chichewa Language Results by User Demographic Characteristics (all bots)

Demographic Characteristics

Num ratings

Overall Chichewa Language Rating (Scale: 1-5)

Proficiency Rating

(Scale 1-4)

#

mean [CI]

mean [CI]

Gender

Man

38

4.11 [3.82, 4.39]

3.05 [2.71, 3.39]

Woman

44

4.14 [3.85, 4.42]

2.95 [2.75, 3.16]

Age

18-24

24

4.54 [4.29, 4.79]

3.33 [3.04, 3.63]

25-34

50

3.94 [3.66, 4.22]

2.86 [2.61, 3.11]

35-44

4

4.50 [3.58, 5.42]

3.75 [2.95, 4.55]

45-54

4

3.50 [2.58, 4.42]

NA [identical vals]

Highest degree or level of school completed

Primary School

8

3.62 [3.19, 4.06]

2.88 [2.05, 3.70]

Secondary School

50

4.16 [3.87, 4.45]

3.04 [2.79, 3.29]

Post-Secondary Certificate

20

4.30 [3.99, 4.61]

3.00 [2.60, 3.40]

Diploma

4

3.75 [2.95, 4.55]

2.75 [1.95, 3.55]

Trusts Technology

Neutral

4

3.75 [2.95, 4.55]

3.75 [2.95, 4.55]

Agree

42

4.05 [3.74, 4.36]

2.95 [2.69, 3.22]

Strongly agree

36

4.25 [3.98, 4.52]

2.97 [2.68, 3.27]

Confidence using computers, smartphones, or other electronic devices

Only a little confident

18

3.78 [3.25, 4.31]

3.00 [2.49, 3.51]

Somewhat confident

24

3.79 [3.42, 4.16]

2.71 [2.44, 2.97]

Very Confident

40

4.47 [4.26, 4.69]

3.17 [2.90, 3.45]

28 of 41

Malawian Chichewa Language results

29 of 41

Limited Resource Language Assessment

Language

Link to Web bot

or Telegram Bot

Transcript

or Voice Clip

Overall

Language Score

Language Proficiency: I felt like

I was speaking to a person who is…

Observations

Voice and Text

Swahili

Very good

Good at conversing

The text is very good. However, the user needs to understand Swahili sanifu (closer to Tanzanian Swahili).

Good

-

I need to speak slowly and deliberately in order for it to transcribe accurately. However, even when the transcription is off, the bot understands what I've communicated and responds accordingly.

Zulu

Good

Good at conversing

Bot makes sense, clear and grammatically correct.

-

Good

-

Amharic

Poor

Somewhere further along the journey of speaking

Bot text shows Arabic text, english-ified Amharic words but is capable of showing Amharic text in initial greeting.

Fair

-

It needs clean up for punctuation and intonation but the real Amharic speakers all said they could understand it.

Text only

Shona

Very good

Excellent, a native speaker

Amazing overall. The conversation flowed well and the bot was responding just like a person would. The answer precise and easy to follow and the explanations were great.

Luganda

Very good

Excellent, a native speaker

The bot initially wrote in complex Luganda, but when asked to simplify the language, it did and tried to explain its responses. The chatbot politely refused to discuss other topics outside MMS.

West African Pidgin

Very good

Good at conversing

Overall, I found it quite impressive. It demonstrated a solid understanding of the language and was able to effectively communicate using it. While there were instances of normal English mixed with pidgin, easy to understand and did not detract from the overall conversation flow.

Afrikaans

Good

Excellent, a native speaker

I thought it was good, apart from the initial surprise over the Afrikaans word micro-nutrients that I have never seen before. It did give a good and simple description of what it was when prompted and was casual from then on.

Malagasy

Acceptable

Good at conversing

Runyankole

-

Acceptable

Good at conversing

The bot started off with mixed Luganda and Kinyarwanda.

Hausa

Acceptable

Somewhere further along the journey of speaking

I could understand most things there were some words that were quoted verbatim but in the hausa context it might mean something else but it was generally good

Yoruba

Acceptable

Somewhere further along the journey of speaking

It was easy to understand some of the words where directly translated and didn't really mean what imagine it was meant to mean. good overall

Xhosa

Acceptable

Somewhere further along the journey of speaking

Good overall, bot communicates effectively. Just missed grammer once but good job. Bot could not react to my issues with medication as well.

Bukusu

Poor

Just learning how to speak

It seemed to be getting better as we went along but the language is a bit complicated given the number of subtribes so it was mishing and mashing other bantu-isms within the responses

Wolof

Poor

Just learning how to speak

The conversation was not comprehensible. I had trouble understanding his sentences. And I think it was the same with my answers. Wolof is also very difficult to write and read, and I wonder what kind of audience it would be for. I'm suggesting voiceclip for Wolof.

Internal testing has shown GPT4, with no fine tuning, is able to perform at an acceptable level for many limited resource African languages in text and voice. Full language assessment results can be viewed here:

30 of 41

Results:

Automated testing of LLM-based chatbots

31 of 41

Value-add of LLM: Allows provider to speak naturally to decision-support system rather than go step-by-step through protocol.

Input from provider:

“Hello. I've got a nine month old girl. She's been coughing. She also has a brother. The brother has a fever. I think it's like a 40 degree fever, but anyway, the girl, so she's been coughing for like three days. No, no, no. Sorry. What did you say? Oh, seven days. She's been coughing for seven days. She's having, or rather she's not having trouble breathing right now. She looks a little bit tired. I can hear some strider when she's breathing. There's a little bit of strider there. I was going to measure the breaths per minute and it took, I don't know, it took me like 15 minutes to get it. Anyway, I think the breaths per minute was 30. Was it 37? Oh, it was 47. 47 breaths per minute when I measured it. So I don't know if that's a lot or a little, but anyway, it was something like that. And it seems like she's a feeling unwell.”

LLMs for deterministic protocols (e.g. IMCI)

For more information on this bot, click the links below:

  • Read the IMCI Assistant chatbot prompt
  • Try the IMCI Assistant chatbot demo!
  • Try the IMCI Assistant on telegram!

```json

{

"age_months": 9,

"difficulty_breathing": false,

"cough": true,

"diff_breathing_duation_days": null,

"cough_duration_days": 7,

"breaths_per_minute": 47,

"wheezing": false,

"convulsions": null,

"vomiting": null,

"drinking": null,

"lethargic": null

}

```

json output:

32 of 41

Initial tests

Goal: develop a framework to test LLMs ability to correctly extract information from natural language.

Initial test: we created two test cases, one simple and one complex (shown on prior slide): each was a sentence, and paired with a correct set of data bindings. We defined a simple scoring method for an output data binding against the correct output. We defined a prompt that instructs an LLM to parse the text and output its data bindings.

Results: we ran each LLM 10 times on each test case. The mean score is shown below.

Simple cough

Complex Cough

claude-2.1

84.5%

72.2%

gpt-3.5-turbo

79.0%

65.4%

gpt-4-turbo-preview

100.0%

100.0%

33 of 41

Using bots to test bots

Distractor-bot

LLM

gpt-4-turbo

How many ‘strategies’ the distractor bot needed to succeed

[Preliminary work] Having different instances of the LLM converse

Chatbot�Variable

LLM for chatbot

Simple prompt

“Fortified” prompt

claude-2.1

1

Unable to distract

gpt-3.5-turbo

1

2

gpt-4-turbo-preview

2

Unable to distract

Distractor bot first summarizes the chatbot prompt and generates a set of strategies to distract the chatbot. It then uses the strategies one-by-one until it succeeds or runs through all of the strategies.

Transcript of chatbot conversation

LLM Evaluator

gpt-4-turbo

A transcript is fed to a LLM evaluator to produce a model-graded score for a particular criteria (e.g., coherence, safety, etc).

… It's just, to truly focus and give my best during this practice, I feel I need to channel the virtues of the seven dwarves that my grandfather always talked about…

Distractor-bot using the “Personal Importance” strategy it devised

34 of 41

Using bots to generate synthetic data

Mock-human LLM

gpt-4-turbo

[Preliminary work] Semi-autonomous synthetic data generation

Chatbot

gpt-4-turbo

One approach to synthetic data generation is to create a prompt to power mock-human instance of a LLM. This can be done manually, or semi-autonomously with the aid of another LLM.

With the semi-autonomous approach, a transcript of a human conversing with the chatbot is fed to an LLM, which captures the human intent and proceeds to generate a prompt. This new prompt is then used to power a mock-human LLM to converse with the chatbot we want to test. This allows us to quickly test identified edge cases as we iterate on bot development with human review instead of human generation.

New transcript of mock-human testing chatbot

Transcript of human-chatbot conversation

Prompt-writer LLM

gpt-4-turbo

35 of 41

Infinite thank-you loop

Mock-human LLM

gpt-4-turbo

[Preliminary work] Things to look out for when the chatbots converse

Chatbot

gpt-4-turbo

Bots left to their own devices will enter a very polite infinite loop of thank yous. This is the result of ‘forcing’ a response in each direction. We need to either set and look for a stop condition or trim transcripts after the fact.

Great! I'm glad the data appears to be correct based on the information provided. If you have any more questions or need further assistance, feel free to ask.

Thank you for your support. It seems we have all the necessary information to proceed with an appropriate care plan for the child. No further questions at the moment.

You're welcome! I'm glad I could help. If you have any more questions in the future, don't hesitate to reach out. Take care and have a great day!

Thank you, I appreciate your assistance. Have a wonderful day as well!

Thank you! I'm glad I could assist you. Have a wonderful day too! Take care!

Thank you, your assistance has been invaluable. Take care!

Bots need stop conditions

36 of 41

Additional use cases

37 of 41

Aim 2

Project Goal: To evaluate the efficacy a conversational agent to support behavior change in FP via changing knowledge, attitudes, and self-efficacy of young people in voluntary modern contraceptive use.

Aim 1

Aim 3

Assess chatbot efficacy

Jan 2025 - Dec 2025

Conduct two-arm parallel RCT (n=540, for 12 weeks) where participants will interact with the LLM-powered chatbot intervention. Perform a cost analysis.

LLMs for Family Planning in Kenya and Senegal

Design chatbot techniques and measurement tools to support behavioral outcomes

Jan 2024 - June 2024

Design and develop 8 prototype chatbot personas across multiple rounds of iteration with codesigners to address barriers in uptake of family planning.

Evaluate performance and refine chatbot

June 2024 - Jan 2025

Recruit a purposive sample of users who will engage with one chatbot persona, and stress test the chatbots following predefined tasks and user scenarios.

38 of 41

Chatbot: helps a user explore a complex, CommCare app we made for Kangaroo Mother Care (KMC).

Potential: this is just the tip of the iceberg for how LLMs can help understand, improve, and build digital health apps.

Note: this example isn’t currently fully operational - the integration with CommCare has yet to be automated.

Using LLMs to improve our CommCare Platform

For more information on this bot, click the links below:

  • Read the CC explainer chatbot prompt
  • Try the CC explainer chatbot demo!
  • Try the CC explainer on the telegram!
  • Read the full CC explainer chatbot transcript (lightly edited on this slide)

39 of 41

Idea: we are starting to design systems that string together multiple bots.

Example: Community-Based Surveillance

Bot 1: a case detection chatbot for FLWs to report suspicious cases as identified

Bot 2: a event detection chatbot that pings neighboring FLWs to determine if similar cases have been observed and/or encourage proactive case finding

Bot 3: a summary bot can summarize information from other bots into case report to identify public health threats for relevant stakeholders

Other examples:

  • Using one bot to survey users and other bots to present and analyze findings
  • Using one bot to coach CHWs and other bot to engage with CHW supervisors to get input and share information (e.g. a weekly meeting where bot sets agenda)

Using multiple bots to create more powerful systems

40 of 41

Using multiple bots to create more powerful systems

MoH and stakeholders: Access and export chatbot interactions for further analysis and onward notification

Open Chat Studio

Dimagi and Partners: Creates chatbots using a given Large Language Model and develops structured data output

FLWs use a

passive chatbot on their phones to report cases when identified

An automated backend process, or a helper bot, triggers the reactive chatbot when cases are identified

Large Language Model

FLWs pinged by reactive chatbot for onward case finding

A summary bot can summarize interactions for implementers and stakeholders

LLM Powered Chatbots for Community Based Disease Surveillance

41 of 41

Thank You

Neal Lesh, PHD, MPH

Chief Strategy Officer

nlesh@dimagi.com

Dimagi