1 of 26

UMSI Data4Good

Updates & Demo

D4G Team

July 25, 2024

2 of 26

Agenda

  1. Introductions (4 mins)
  2. Products (2 mins)
  3. Hangul Design (12 mins)
  4. Hangul Summary Generation (10 mins)
  5. Other Algorithm Improvements (2 mins)
  6. Hangul Demo (16 mins)
  7. Q&A (12 mins)
  8. Next Steps (2 mins)

3 of 26

Introductions

4 of 26

Intro

The Data4Good center is working on projects that are at the center of the UMSI mission: “We create and share knowledge so that people will use information — with technology — to build a better world.”

Our History:

The UMSI Data4Good (D4G) center brings together data from nonprofit organizations into larger datasets for benchmarking and trend analysis. By empowering nonprofits with comprehensive data, we contribute to a broader understanding of development and relief programs, ultimately enhancing decision-making in nonprofit work.

Our Values:

  • Transparency and equal access to information.
  • Utilizing data analysis to positively impact lives through nonprofit organizations.

Thank you for joining us. We look forward to your feedback and insights.

5 of 26

Team Members

6 of 26

Products

7 of 26

Data4Good Products

  • Chetah: a search engine for nonprofits reports published on the web. It locates reports with a state of the art deep learning algorithm.

  • Hangul: An NLP-based assistant designed for digital curators at ReliefWeb to handle a larger volume of documents.

8 of 26

Hangul Design

9 of 26

Hangul Timeline

W2023

  • Demo of Hangul 1.0 to ReliefWeb team (Mar. 7)
  • UMSI Poster Presentation (Apr. 17)
  • Added Hangul to the D4G website

S2023

  • Hangul 2.0 work
  • Added models for Theme, Language, Disaster Type, Locations, Improved Summary, Markdown text
  • Added the front-end and back-end architecture

F2023

  • Hangul 2.0 tuning
  • Presentation to systems advisor, Intelligems CTO (Oct. 18)
  • Summary experiments

S2024

  • Hangul 2.1 presentation
  • Demo to ReliefWeb team (July 25)

W2024

  • Hangul 2.1
  • Changed back-end structure; process PDF in stages
  • PythonAnywhere memory constraint
  • Algorithm testing, especially for the Summary

10 of 26

Hangul 2.1: What Changed?

Result Item

Hangul 1.0

Hangul 2.0

Filename

Report Title

✔+

Report Author

Document Creation Date

Document Modified Date

Report Type

Number of Pages

Result Item

Hangul 1.0

Hangul 2.0

Document Language

Document Locations

✔+

Disaster Type

Themes

Markdown

Keywords

Summary

Result Item

Hangul 2.1

Memory Re Architecture

11 of 26

Hangul 2.1: Application Development

Overall Updates

  • Implemented new theme

  • Updated mobile-view

UI Updates

  • Cleaner interface

  • Feedback to user on backend progress

Front End Updates

  • Improved system robustness�
  • Admin metric alerts

12 of 26

Hangul 2.1 Data Structure

Objective:

    • Enhance the quality of metadata features and generate new ones.

Workflow:

    • Retrieve report data using the ReliefWeb API.
    • Parse stored reports and extract various related metadata features.
    • Analyze and process data using NLP techniques to create a dataset suitable for training and testing models.

13 of 26

Hangul API

Memory optimization

Earlier in the year we encounter memory limitations, that required us to run Hangul under 3 GB.

We overcame this limitation by reengineering our API for a two-stage document processing.

14 of 26

Hangul 2.1: Improved API Handling

Hangul

Python Anywhere

HTTP Request

Bad response?

Less than 3 tries?

2^n + rand delay

YES

NO

HTTP Response

YES

15 of 26

Hangul Summary

16 of 26

Summary: Generation

17 of 26

Summary: Quality Metrics

BERTScore is an evaluation metric that leverages BERT embeddings to measure the semantic similarity between generated and reference texts. Instead of relying on exact word matches, BERTScore captures the contextual and nuanced meanings of words in their specific contexts.

Example:

Reference: "The cat sat on the mat."

Generated: "The feline rested on the rug."

Example Scores

18 of 26

Summary Generation Time

19 of 26

Summary Generation Time

20 of 26

Summary Generation Time

21 of 26

Other Algorithm Improvements

22 of 26

Hangul 2.1: Success Metrics

23 of 26

Hangul Demo

24 of 26

Questions

Your Questions For Us

Our Questions For You

  • On average, how long does it take an editor to complete a PDF
  • Given the context of assistant-AI, what would be the obstacles to adoption?
  • Given what we have tried, what would you do next?
  • How can we make this better for the end-user?
  • As a user what would you expect to see and be able to do on the D4G website?
  • If you were to present this to a nonprofit leader, what would you present?
  • What constitutes “World” for location?

25 of 26

Next Steps

  • Continued summary improvement
  • Determine the factors driving the time-to-analyze
  • Share APIs between Hangul and Chetah
  • Have Hangul results feed the Chetah dataset
  • Develop an MVP for ReliefChat

26 of 26

Thank you!