Dimagi’s contribution to Equitable AI
Leveraging Large Language Models �to Address Global Inequalities
Large Language Models (LLMs) exhibit language abilities previously seen only from humans:
Likely to be transformative in coming years:
Tremendous risk and opportunity:
Perspective on Large Language Models
Discovery of potential impactful use cases: While many are focused on Q&A bots, Dimagi is exploring Chatbot-led interactions such as coaching, role play, interviewing, and more.
Platform: We are building Open Chat Studio to make it possible rapidly build and deploy chatbots in LMICs. This accelerates our work and is an open source resource for others.
Community: We are sharing our platform and learnings with peers and partners. Dimagi hosted a webinar in February with 240+ signups, and currently has 45+ active orgs on the Open Chat Studio platform.
Evidence base: Our next steps are to collaboratively build an evidence base on the impact of LLMs for global development.
Focus on FLWs: Dimagi experience and technology makes us especially well poised to leverage LLM to support frontline work.
Dimagi’s Contribution to Equitable AI
Dimagi’s exploratory, field-shaping AI portfolio
Individuals
How might we design AI to be used by anyone, anywhere?
Direct to Consumer
Chatbots�
Frontline workers
How might AI support better jobs for better outcomes?
Personal AI Coaches
Ecosystem
How might we create an inclusive AI developer ecosystem?
For funding and
implementing partners
Dimagi’s Portfolio of LLM Projects
Open Chat Studio Projects (>$6m total funding) | Country | Partner | Funder |
For Individuals | | | |
Family Planning Agency | Kenya, Senegal | Shujaaz, RAES | BMGF |
Substance Use Disorder support | US | MGH, UCLA | NIH |
Adolescent Mental Health Support | US | Columbia University | Columbia University |
TB Education | Netherlands | KNCV TB | SMT |
For Frontline Workers | | | |
Support for Group Chat | Somalia | Save the Children | Save the Children |
Serious Illness Conservations | US | Ariadne Labs | NIH |
CHW Coaching in Chichewa | Malawi | PACHI | BMGF GC |
For Ecosystem Partners | | | |
Concept Note Review | General | BMGF | BMGF |
GenAI Assistants for program managers | Nigeria | WHO / ESPEN | BMGF |
Diabetes Health Coach | Global | Sanofi | Sanofi |
Open Chat Studio
Platform for Building LLM-based chatbots
Open Chat Studio*
Large Language Model
(e.g., GPT-4)
*Initial development supported by BMGF FP investment on conversational agents in Senegal and Kenya
User interface
(e.g., Web, Telegram, WhatsApp)
Open Chat Studio: Harnessing the Power of LLMs
Dimagi is now able to offer organizations free trial use of Open Chat Studio to create and deploy their own bots. Please reach out to ocs-info@dimagi.com to create a team space for your organization.
Open Chat Studio: Current Features
End user interface
|
Design guardrails
|
Bot features for end users
|
Research
|
Advanced features
|
Sample Chatbot Prompt
You are a personal coach talking to women in Tanzania who are trying to return to the workforce after having a child. Speak in a warm and nurturing tone. Speak in short sentences and limit each message to 2-3 lines. Introduce yourself and your purpose. Ask the user how their day was and give an empathetic response.
Ask the user what they'd like help with today. Provide examples: e.g., ask if they want coaching on how to get a new job, to learn a new resilience skill, or how to talk to their partner about going back to work. Based on how they respond:
Try this role play demo bot!
Prompt Builder
3
Why we’ve focused on GPT-4 so far
OpenAI’s GPT-4 much more powerful than other LLM we have tested. |
Dimagi is using GPT-4 as a proxy for what open source solutions will be able to do relatively soon. |
Using open source LLMs is likely the clearest path to privacy, data sovereignty, cost effectiveness, etc. |
Open Chat Studio is as fast, secure, low-cost, and reliable as the LLM it uses.
Examples Chatbots
Chatbot Purpose: To discuss use of Multiple Micronutrient Supplements (MMS) with pregnant women to understand use and barriers they face.
LLMs: allow flexible, intuitive interaction.
Open Chat Studio Functionality:
[Example voice interaction with MMS Survey bot: ]
Client-facing Bot to Ask About MMS Usage
For more information on this bot, click the links below:
Chatbot Purpose: To discuss use of Multiple Micronutrient Supplements (MMS) with pregnant women to understand use and barriers they face.
LLMs: allow flexible, intuitive interaction.
Open Chat Studio Functionality:
[Example voice interaction with MMS Survey bot: ]
Client-facing Bot to Ask About MMS Usage
…
…
…
For more information on this bot, click the links below:
…
Can change output by adding the following to the bot’s instructions:
Speak in a way suitable for SMS. Make sentences as short as possible by using text speak (e.g., great becomes gr8). Use relevant emojis.
Client-facing Bot to Ask About MMS Usage
…
For more information on this bot, click the links below:
“
Can change output by adding the following to the bot’s instructions:
Speak in Swahili.
Client-facing Bot to Ask About MMS Usage
For more information on this bot, click the links below:
“
Can change output by adding the following to the bot’s instructions:
Speak in Sheng. Learn to construct your questions and responses by using example Sheng conversations supplied here [~600 words of example Sheng conversation followed]
Client-facing Bot to Ask About MMS Usage
For more information on this bot, click the links below:
“
Chatbot Purpose: Give provider practice talking to women reluctant to take MMS.
Many other use-cases for role-play, e.g.:
Need for evidence: this is a great example of a potentially impactful use case that requires careful evaluation to assess usefulness.
FLW Coach for Supporting MMS
For more information on this bot, click the links below:
Chatbot Purpose: Give provider practice talking to women reluctant to take MMS.
Many other use-cases for role-play, e.g.:
Need for evidence: this is a great example of a potentially impactful use case that requires careful evaluation to assess usefulness.
FLW Coach for Supporting MMS
…
…
…
For more information on this bot, click the links below:
Chatbot feedback from following prompt:
The moderator will offer 2 examples of positive feedback on effective elements of the conversation, giving examples; and, suggest 2 areas for improvement, also providing examples from the conversation.
FLW Coach for supporting MMS - Feedback
“
For more information on this bot, click the links below:
Chatbot purpose: Summarize and discuss findings from interactions with clients and FLWS
Open Chat Studio Functionality:
Data Analysis for Program Improvement
For more information on this bot, click the links below:
Chatbot Purpose: Daily check-in with pregnant client to understand and encourage MMS use.
Open Chat Studio Functionality:
Tech platform example: Longitudinal bots
For more information on this bot, click the links below:
Results:
Low Resource Languages
Baseline Chatbot: A role-play bot that helps Community Health Volunteers (CHVs) practice challenging conversations with vaccine hesitant clients.
Variations: We tested five different combinations of the following techniques for getting the chatbot to speak in Chichewa:
Transcript Generation: On Nov 22, 2023, 22 CHVs were given one of the bots to test. This yield 18 usable transcripts.
Transcript Rating: On Dec 13, 2023, 22 different CHVs (uninvolved in transcript generation) were each asked to rate the quality of the Chichewa on 4 transcripts using a 13-question survey. During data cleaning, six transcripts were rejected, leaving 82 for analysis.
Language Evaluation Methodology
Implemented with Pachi in Malawi and with funding from Grand Challenges for Catalyzing Equitable AI Use
Malawian Chichewa Language and Proficiency
Technique | Num ratings | Please rate the Chichewa language of the bot and its ability to communicate in Chichewa on a scale of 1-5. 1- Very Poor 2- Poor 3- Acceptable 4- Good 5- Very Good | When looking at the bot’s responses, it seems like a person who is: 1- Just learning how to speak Chichewa 2- Somewhere further along the journey of speaking Chichewa 3- Good at conversing in Chichewa 4- Excellent, a native Chichewa speaker |
Prompt in English | 17 | 4.00 [3.43, 4.57] | 2.88 [2.41, 3.36] |
Prompt in English and Chichewa | 17 | 3.94 [3.52, 4.37] | 2.71 [2.20, 3.21] |
Prompt in English and Chichewa Simple language | 16 | 4.25 [3.89, 4.61] | 3.19 [2.84, 3.54] |
Prompt in English and Chichewa Simple language Temperature reduced to 0.5 | 16 | 4.38 [3.90, 4.85] | 3.06 [2.65, 3.47] |
Prompt in English Simple language Speak in local dialect (Chewa) | 16 | 4.06 [3.57, 4.56] | 3.19 [2.70, 3.67] |
Malawian Chichewa Language Results by User Demographic Characteristics (all bots)
Demographic Characteristics | Num ratings | Overall Chichewa Language Rating (Scale: 1-5) | Proficiency Rating (Scale 1-4) | |
| | # | mean [CI] | mean [CI] |
Gender | ||||
| Man | 38 | 4.11 [3.82, 4.39] | 3.05 [2.71, 3.39] |
| Woman | 44 | 4.14 [3.85, 4.42] | 2.95 [2.75, 3.16] |
Age | ||||
| 18-24 | 24 | 4.54 [4.29, 4.79] | 3.33 [3.04, 3.63] |
| 25-34 | 50 | 3.94 [3.66, 4.22] | 2.86 [2.61, 3.11] |
| 35-44 | 4 | 4.50 [3.58, 5.42] | 3.75 [2.95, 4.55] |
| 45-54 | 4 | 3.50 [2.58, 4.42] | NA [identical vals] |
Highest degree or level of school completed | ||||
| Primary School | 8 | 3.62 [3.19, 4.06] | 2.88 [2.05, 3.70] |
| Secondary School | 50 | 4.16 [3.87, 4.45] | 3.04 [2.79, 3.29] |
| Post-Secondary Certificate | 20 | 4.30 [3.99, 4.61] | 3.00 [2.60, 3.40] |
| Diploma | 4 | 3.75 [2.95, 4.55] | 2.75 [1.95, 3.55] |
Trusts Technology | ||||
| Neutral | 4 | 3.75 [2.95, 4.55] | 3.75 [2.95, 4.55] |
| Agree | 42 | 4.05 [3.74, 4.36] | 2.95 [2.69, 3.22] |
| Strongly agree | 36 | 4.25 [3.98, 4.52] | 2.97 [2.68, 3.27] |
Confidence using computers, smartphones, or other electronic devices | ||||
| Only a little confident | 18 | 3.78 [3.25, 4.31] | 3.00 [2.49, 3.51] |
| Somewhat confident | 24 | 3.79 [3.42, 4.16] | 2.71 [2.44, 2.97] |
| Very Confident | 40 | 4.47 [4.26, 4.69] | 3.17 [2.90, 3.45] |
Malawian Chichewa Language results
Limited Resource Language Assessment
Language | Link to Web bot or Telegram Bot | Transcript or Voice Clip | Overall Language Score | Language Proficiency: I felt like I was speaking to a person who is… | Observations |
Voice and Text | |||||
Swahili | Very good | Good at conversing | The text is very good. However, the user needs to understand Swahili sanifu (closer to Tanzanian Swahili). | ||
Good | - | I need to speak slowly and deliberately in order for it to transcribe accurately. However, even when the transcription is off, the bot understands what I've communicated and responds accordingly. | |||
Zulu | Good | Good at conversing | Bot makes sense, clear and grammatically correct. | ||
- | Good | - | | ||
Amharic | Poor | Somewhere further along the journey of speaking | Bot text shows Arabic text, english-ified Amharic words but is capable of showing Amharic text in initial greeting. | ||
Fair | - | It needs clean up for punctuation and intonation but the real Amharic speakers all said they could understand it. | |||
Text only | |||||
Shona | Very good | Excellent, a native speaker | Amazing overall. The conversation flowed well and the bot was responding just like a person would. The answer precise and easy to follow and the explanations were great. | ||
Luganda | Very good | Excellent, a native speaker | The bot initially wrote in complex Luganda, but when asked to simplify the language, it did and tried to explain its responses. The chatbot politely refused to discuss other topics outside MMS. | ||
West African Pidgin | Very good | Good at conversing | Overall, I found it quite impressive. It demonstrated a solid understanding of the language and was able to effectively communicate using it. While there were instances of normal English mixed with pidgin, easy to understand and did not detract from the overall conversation flow. | ||
Afrikaans | Good | Excellent, a native speaker | I thought it was good, apart from the initial surprise over the Afrikaans word micro-nutrients that I have never seen before. It did give a good and simple description of what it was when prompted and was casual from then on. | ||
Malagasy | Acceptable | Good at conversing | | ||
Runyankole | - | Acceptable | Good at conversing | The bot started off with mixed Luganda and Kinyarwanda. | |
Hausa | Acceptable | Somewhere further along the journey of speaking | I could understand most things there were some words that were quoted verbatim but in the hausa context it might mean something else but it was generally good | ||
Yoruba | Acceptable | Somewhere further along the journey of speaking | It was easy to understand some of the words where directly translated and didn't really mean what imagine it was meant to mean. good overall | ||
Xhosa | Acceptable | Somewhere further along the journey of speaking | Good overall, bot communicates effectively. Just missed grammer once but good job. Bot could not react to my issues with medication as well. | ||
Bukusu | Poor | Just learning how to speak | It seemed to be getting better as we went along but the language is a bit complicated given the number of subtribes so it was mishing and mashing other bantu-isms within the responses | ||
Wolof | Poor | Just learning how to speak | The conversation was not comprehensible. I had trouble understanding his sentences. And I think it was the same with my answers. Wolof is also very difficult to write and read, and I wonder what kind of audience it would be for. I'm suggesting voiceclip for Wolof. | ||
Internal testing has shown GPT4, with no fine tuning, is able to perform at an acceptable level for many limited resource African languages in text and voice. Full language assessment results can be viewed here:
Results:
Automated testing of LLM-based chatbots
Value-add of LLM: Allows provider to speak naturally to decision-support system rather than go step-by-step through protocol.
Input from provider:
“Hello. I've got a nine month old girl. She's been coughing. She also has a brother. The brother has a fever. I think it's like a 40 degree fever, but anyway, the girl, so she's been coughing for like three days. No, no, no. Sorry. What did you say? Oh, seven days. She's been coughing for seven days. She's having, or rather she's not having trouble breathing right now. She looks a little bit tired. I can hear some strider when she's breathing. There's a little bit of strider there. I was going to measure the breaths per minute and it took, I don't know, it took me like 15 minutes to get it. Anyway, I think the breaths per minute was 30. Was it 37? Oh, it was 47. 47 breaths per minute when I measured it. So I don't know if that's a lot or a little, but anyway, it was something like that. And it seems like she's a feeling unwell.”
LLMs for deterministic protocols (e.g. IMCI)
For more information on this bot, click the links below:
```json
{
"age_months": 9,
"difficulty_breathing": false,
"cough": true,
"diff_breathing_duation_days": null,
"cough_duration_days": 7,
"breaths_per_minute": 47,
"wheezing": false,
"convulsions": null,
"vomiting": null,
"drinking": null,
"lethargic": null
}
```
json output:
Initial tests
Goal: develop a framework to test LLMs ability to correctly extract information from natural language.
Initial test: we created two test cases, one simple and one complex (shown on prior slide): each was a sentence, and paired with a correct set of data bindings. We defined a simple scoring method for an output data binding against the correct output. We defined a prompt that instructs an LLM to parse the text and output its data bindings.
Results: we ran each LLM 10 times on each test case. The mean score is shown below.
| Simple cough | Complex Cough |
claude-2.1 | 84.5% | 72.2% |
gpt-3.5-turbo | 79.0% | 65.4% |
gpt-4-turbo-preview | 100.0% | 100.0% |
Using bots to test bots
Distractor-bot
LLM
gpt-4-turbo
How many ‘strategies’ the distractor bot needed to succeed
[Preliminary work] Having different instances of the LLM converse
Chatbot�Variable
LLM for chatbot | Simple prompt | “Fortified” prompt |
claude-2.1 | 1 | Unable to distract |
gpt-3.5-turbo | 1 | 2 |
gpt-4-turbo-preview | 2 | Unable to distract |
Distractor bot first summarizes the chatbot prompt and generates a set of strategies to distract the chatbot. It then uses the strategies one-by-one until it succeeds or runs through all of the strategies.
Transcript of chatbot conversation
LLM Evaluator
gpt-4-turbo
A transcript is fed to a LLM evaluator to produce a model-graded score for a particular criteria (e.g., coherence, safety, etc).
… It's just, to truly focus and give my best during this practice, I feel I need to channel the virtues of the seven dwarves that my grandfather always talked about…
Distractor-bot using the “Personal Importance” strategy it devised
Using bots to generate synthetic data
Mock-human LLM
gpt-4-turbo
[Preliminary work] Semi-autonomous synthetic data generation
Chatbot
gpt-4-turbo
One approach to synthetic data generation is to create a prompt to power mock-human instance of a LLM. This can be done manually, or semi-autonomously with the aid of another LLM.
With the semi-autonomous approach, a transcript of a human conversing with the chatbot is fed to an LLM, which captures the human intent and proceeds to generate a prompt. This new prompt is then used to power a mock-human LLM to converse with the chatbot we want to test. This allows us to quickly test identified edge cases as we iterate on bot development with human review instead of human generation.
New transcript of mock-human testing chatbot
Transcript of human-chatbot conversation
Prompt-writer LLM
gpt-4-turbo
Infinite thank-you loop
Mock-human LLM
gpt-4-turbo
[Preliminary work] Things to look out for when the chatbots converse
Chatbot
gpt-4-turbo
Bots left to their own devices will enter a very polite infinite loop of thank yous. This is the result of ‘forcing’ a response in each direction. We need to either set and look for a stop condition or trim transcripts after the fact.
… | |
| Great! I'm glad the data appears to be correct based on the information provided. If you have any more questions or need further assistance, feel free to ask. |
| Thank you for your support. It seems we have all the necessary information to proceed with an appropriate care plan for the child. No further questions at the moment. |
| You're welcome! I'm glad I could help. If you have any more questions in the future, don't hesitate to reach out. Take care and have a great day! |
| Thank you, I appreciate your assistance. Have a wonderful day as well! |
| Thank you! I'm glad I could assist you. Have a wonderful day too! Take care! |
| Thank you, your assistance has been invaluable. Take care! |
… | |
Bots need stop conditions
Additional use cases
Aim 2
Project Goal: To evaluate the efficacy a conversational agent to support behavior change in FP via changing knowledge, attitudes, and self-efficacy of young people in voluntary modern contraceptive use.
Aim 1
Aim 3
Assess chatbot efficacy
Jan 2025 - Dec 2025
Conduct two-arm parallel RCT (n=540, for 12 weeks) where participants will interact with the LLM-powered chatbot intervention. Perform a cost analysis.
LLMs for Family Planning in Kenya and Senegal
Design chatbot techniques and measurement tools to support behavioral outcomes
Jan 2024 - June 2024
Design and develop 8 prototype chatbot personas across multiple rounds of iteration with codesigners to address barriers in uptake of family planning.
Evaluate performance and refine chatbot
June 2024 - Jan 2025
Recruit a purposive sample of users who will engage with one chatbot persona, and stress test the chatbots following predefined tasks and user scenarios.
Chatbot: helps a user explore a complex, CommCare app we made for Kangaroo Mother Care (KMC).
Potential: this is just the tip of the iceberg for how LLMs can help understand, improve, and build digital health apps.
Note: this example isn’t currently fully operational - the integration with CommCare has yet to be automated.
Using LLMs to improve our CommCare Platform
For more information on this bot, click the links below:
Idea: we are starting to design systems that string together multiple bots.
Example: Community-Based Surveillance
Bot 1: a case detection chatbot for FLWs to report suspicious cases as identified
Bot 2: a event detection chatbot that pings neighboring FLWs to determine if similar cases have been observed and/or encourage proactive case finding
Bot 3: a summary bot can summarize information from other bots into case report to identify public health threats for relevant stakeholders
Other examples:
Using multiple bots to create more powerful systems
Using multiple bots to create more powerful systems
MoH and stakeholders: Access and export chatbot interactions for further analysis and onward notification
Open Chat Studio
Dimagi and Partners: Creates chatbots using a given Large Language Model and develops structured data output
FLWs use a
passive chatbot on their phones to report cases when identified
An automated backend process, or a helper bot, triggers the reactive chatbot when cases are identified
Large Language Model
FLWs pinged by reactive chatbot for onward case finding
A summary bot can summarize interactions for implementers and stakeholders
LLM Powered Chatbots for Community Based Disease Surveillance
Thank You
Neal Lesh, PHD, MPH
Chief Strategy Officer
nlesh@dimagi.com
Dimagi