The mission of our collaborative initiative is to democratize access to AI, ensuring equitable learning opportunities for all children and educators, regardless of their socioeconomic status.
Evidence-based & technically sound
Equitable and affordable
Implementable
and able to improve
learning at scale
Our North Star:
We want to support AI tools for low-and-middle income countries (LMICs) to be:
"We shape the world’s best technologies to transform education systems and help those learning the least"
We believe AI has the potential to radically transform education systems and improve learning.
SCALABLE QUALITY ASSURANCE SYSTEMS
REAL-WORLD INNOVATION AND SCALING
We are doubling down on two areas of high need/demand
Develop the solution / product
Pedagogical foundations
Product and AI evaluation
Safety and ethics
▪ ToC’s alignment to evidence / science
of learning
▪ Content quality
▪ Child safety / do no harm
▪ Data security and governance
▪ AI model benchmarks
▪ Evals on outputs
▪ Tech stack
▪ UX design
Get it in the hands of users
(under the right conditions)
RTM / access
Product uptake
Implementation
▪ B2C / B2B / B2G
▪ Infrastructure & device access
▪ Product cost & affordability
▪ Fidelity of implementation
▪ Supervision of product use
▪ User registration & trial
Experience & perceptions
Engagement
Understand the user experience
▪ Intuitive usability
▪ Satisfaction (NPS)
▪ Trust
▪ Dosage / time on task
▪ Engagement quality / use as intended
▪ Drop off rate
Intermediate user changes
Impact
(on outcome of interest)
▪ Cognitive
▪ Affective
▪ Behavioural
▪ Learning gains & subject mastery
▪ Eduficiency (teaching efficiency & productivity)
Cost-effectiveness
Measure the impact
▪ Impact per dollar spent (fixed / variable)
Quality assurance is central to our digital public good offer
1
2
4
3
Develop the solution / product
Pedagogical foundations
Product and AI evaluation
Safety and ethics
▪ ToC’s alignment to evidence / science
of learning
▪ Content quality
▪ Child safety / do no harm
▪ Data security and governance
▪ AI model benchmarks
▪ Evals on outputs
▪ Tech stack
▪ UX design
Get it in the hands of users
(under the right conditions)
RTM / access
Product uptake
Implementation
▪ B2C / B2B / B2G
▪ Infrastructure & device access
▪ Product cost & affordability
▪ Fidelity of implementation
▪ Supervision of product use
▪ User registration & trial
Experience & perceptions
Engagement
Understand the user experience
▪ Intuitive usability
▪ Satisfaction (NPS)
▪ Trust
▪ Dosage / time on task
▪ Engagement quality / use as intended
▪ Drop off rate
Intermediate user changes
Impact
(on outcome of interest)
▪ Cognitive
▪ Affective
▪ Behavioural
▪ Learning gains & subject mastery
▪ Eduficiency (teaching efficiency & productivity)
Cost-effectiveness
Measure the impact
▪ Impact per dollar spent (fixed / variable)
We support with content curation, benchmarks and evals to do controlled environment testing of AI tools
1
2
4
3
Developed three benchmarks: Pedagogy; SEND; visual reasoning.
Plus specific task evals: Gender balance; phonics alignment; EMIS data retrieval.
Plus agentic model for tracking products, including data privacy and security.
Creating a digital public good library of high-quality content for training/fine-tune/RAG: includes scalable evals for quality of textbooks; lesson-plans; story-books.
We’re developing benchmarks for educational use cases and filling key gaps—such as for Pedagogical knowledge
General reasoning | 1 | General reasoning |
Pedagogy | 2.1 | Pedagogical knowledge |
2.2 | Pedagogy of generated outputs | |
2.3 | Pedagogical interactions | |
Educational content | 3.1 | Content knowledge |
3.2 | Content alignment | |
Assessment | 4.1 | Scoring and grading |
4.2 | Feedback with reasoning | |
Ethics and bias | 5 | Ethics and bias |
Digitisation / accessibility | 6.1 | Multimodal capabilities |
6.2 | Multilingual capabilities |
Performance of LLMs on pedagogical knowledge vs. cost
See our website for a complete / up to date list of LLMs assessed
Average human teacher score in Chile (50%)
KEY BENCHMARK NEEDS FOR AI-EDTECH
OUR PEDAGOGICAL KNOWLEDGE BENCHMARK
We’re identifying gaps for education which the AI community overlooks
General reasoning | 1 | General reasoning |
Pedagogy | 2.1 | Pedagogical knowledge |
2.2 | Pedagogy of generated outputs | |
2.3 | Pedagogical interactions | |
Educational content | 3.1 | Content knowledge |
3.2 | Content alignment | |
Assessment | 4.1 | Scoring and grading |
4.2 | Feedback with reasoning | |
Ethics and bias | 5 | Ethics and bias |
Digitisation / accessibility | 6.1 | Multimodal capabilities |
6.2 | Multilingual capabilities |
Grade 7 Zambian visual reasoning exams
Model | Accuracy |
Gemini-2.5-pro | 59% |
GPT-5 Medium thinking time | 52% |
GPT-4o | 34% |
This matters, as early grade maths is learned through pictures.
KEY BENCHMARK NEEDS FOR AI-EDTECH
THE AI WORK IS CELEBRATING AI MODELS BEATING TOP MATH PROBLEMS—WHILE WE FIND THEY STRUGGLE WITH GRADE 7 ARICA VISUAL EXAMS
Develop the solution / product
Pedagogical foundations
Product and AI evaluation
Safety and ethics
▪ ToC’s alignment to evidence / science
of learning
▪ Content quality
▪ Child safety / do no harm
▪ Data security and governance
▪ AI model benchmarks
▪ Evals on outputs
▪ Tech stack
▪ UX design
Get it in the hands of users
(under the right conditions)
RTM / access
Product uptake
Implementation
▪ B2C / B2B / B2G
▪ Infrastructure & device access
▪ Product cost & affordability
▪ Fidelity of implementation
▪ Supervision of product use
▪ User registration & trial
Experience & perceptions
Engagement
Understand the user experience
▪ Intuitive usability
▪ Satisfaction (NPS)
▪ Trust
▪ Dosage / time on task
▪ Engagement quality / use as intended
▪ Drop off rate
Intermediate user changes
Impact
(on outcome of interest)
▪ Cognitive
▪ Affective
▪ Behavioural
▪ Learning gains & subject mastery
▪ Eduficiency (teaching efficiency & productivity)
Cost-effectiveness
Measure the impact
▪ Impact per dollar spent (fixed / variable)
We support innovators with small grants where we think there’s big public upside
1
2
4
3
Our learning by doing grants support targeted R&D
Plus support on evidence – and moving to offer support on evals.
To date small grants (<100k) focussed on digital public good components.
These projects have collectively reached millions of children with their innovations.
Plus we are investing in improving the training corpus including Voice-AI (in Tanzania).
Develop the solution / product
Pedagogical foundations
Product and AI evaluation
Safety and ethics
▪ ToC’s alignment to evidence / science
of learning
▪ Content quality
▪ Child safety / do no harm
▪ Data security and governance
▪ AI model benchmarks
▪ Evals on outputs
▪ Tech stack
▪ UX design
Get it in the hands of users
(under the right conditions)
RTM / access
Product uptake
Implementation
▪ B2C / B2B / B2G
▪ Infrastructure & device access
▪ Product cost & affordability
▪ Fidelity of implementation
▪ Supervision of product use
▪ User registration & trial
Experience & perceptions
Engagement
Understand the user experience
▪ Intuitive usability
▪ Satisfaction (NPS)
▪ Trust
▪ Dosage / time on task
▪ Engagement quality / use as intended
▪ Drop off rate
Intermediate user changes
Impact
(on outcome of interest)
▪ Cognitive
▪ Affective
▪ Behavioural
▪ Learning gains & subject mastery
▪ Eduficiency (teaching efficiency & productivity)
Cost-effectiveness
Measure the impact
▪ Impact per dollar spent (fixed / variable)
We are reimagining rapid efficacy studies to improve evidence at scale
1
2
4
3
We will manage a fund for rapid studies of real-world impact – initial project in SL. – expanding to full fund.
Have curated a set of shared, measurable indicators that can scale at lower costs, including learning outcomes.
We create evidence on what works
We curate what is out there
Track emerging evidence on AI tools for education.
Developed multi-agent systems to track products and find the public evidence claims.
(Beta) scaling AI-quality judgements against the ESSA frameworks
Developing web-app for Ministries/funders to check quality claims.
We are designing two district level projects to explore what is needed to take ideas into practice
Two countries, one county/ district each ⟶ reaching ~800 schools, 500,000+ learners impacted.
(Kenya: 281,973, Sierra Leone: 254,639)
WITH A FEATURE PHONE
A voice AI tutor
We will test how can AI support teaching and learning in a systematic way
A STUDENT
SMS practice exercises
WITH A SMART PHONE OR TABLET
Learning platforms
A TEACHER
A QUALITY ASSURANCE OFFICER
Automatic marking
WITH A SMART PHONE OR TABLET
Coaching & feedback
Lesson planning
Tailored guidance for each school
Legend:
Underlying tech is not strong yet, but will be soon
Requires fully integrated solutions between school / home / EMIS
WITH A SMART PHONE OR TABLET
Tailored support to teachers
School prioritization
AI Trainings
Training Space
GenAI NGO Peer Learning Cohort
Africa ICT Right
Dignitas
EducAid Sierra Leone
Orkeeswa, Inc.
Phoenix Foundation
ReachAll
Rising Hope Girls Educational Foundation
The Agahozo-Shalom Youth Village
Yiya Engineering Solutions
Zimbabwe Information and Technology Empowerment Trust
team4tech.org | 14