1 of 25

Extracting Image Entities with Machine Learning

Alexander O’Neill (UPEI) - 2022-08-05

IslandoraCon 2022

2 of 25

Chronicling America

3 of 25

Newspaper Navigator

  • Developed by Ben Lee, the “Innovator in Residence” at the Library of Congress in 2020
  • “Newspaper Navigator” was a 3 phase project
    • Phase 1: Extract images from scanned newspaper pages using Machine Learning
    • This is what our project is based on
    • The results of further phases can be seen at the LoC Chronicling America Newspaper Navigator website.

4 of 25

Learning Model

  • We use LoC’s learning model directly
  • The “Beyond Words” crowd-sourcing project
  • Trained on bounding box data of
    • Photographs
    • Headlines
    • Comics
    • Advertisements
    • Maps

5 of 25

CS 4820

  • 4th year Computer Science course at UPEI taught by Dr. David LeBlanc
  • “Clients” from other UPEI schools and elsewhere pitch projects
    • Then act as client. Students plan and manage the project themselves.
  • Students assemble in to small teams
  • Robertson Library has made good use of work from these projects in the past
    • A rules engine for Fedora 3
    • TinyMCE enhancements for annotations

6 of 25

The Team

  • Client: UPEI Robertson Library
    • Alexander O’Neill, Rosie Le Faive
  • CS 4810 Project Team:
    • Kosay Jabre - Team lead
    • Alexander Cairns
    • Remah Badr
    • Immanuel Olisa
    • Jiarui Qu

7 of 25

Project Components

  • Image Segmentation API
  • PHP client library
  • Islandora 7 Solution Pack

8 of 25

Image Segmentation API

9 of 25

Image Segmentation API

  • Uses the training data released by the Newspaper Navigator project
  • The original program was wrapped in a Docker container
    • Built around AWS-based servers

10 of 25

  • The team used the training data but built their own Python application
  • API endpoint with FastAPI
  • Detectron
    • Made by Facebook, experimental framework for image detection
  • PyTorch
    • Machine learning framework

11 of 25

Processing Pipeline

Resizing

  • IslandNewspapers scanned images are very high-res
  • Without resizing, the extractor found lots of false positives
    • Specks of ink, periods, text decorations interpreted as images

Source: “CS 4820 Final Report”, Jabre et. al., 2021

12 of 25

Processing Pipeline

  • The detected images’ coordinates are scaled back up to the original image
  • Then that image is cropped out of the original at full resolution

Normalize output bounding boxes

Source: “CS 4820 Final Report”, Jabre et. al., 2021

13 of 25

Processing Pipeline

  • The response is a JSON file with
    • The extracted images as download URLs
    • Embedded OCR and hOCR

Create HTTP response

Source: “CS 4820 Final Report”, Jabre et. al., 2021

14 of 25

GPU Support

  • Nvidia CUDA support can be enabled with a configuration flag
  • Observed a performance improvement of about 150%
  • Drivers are included optionally in the downloadable Docker image
    • AWS has machines with graphics cards included

15 of 25

PHP Client Library

16 of 25

  • PHP class to represent API responses
  • Pure PHP, no need to port for Drupal 8+ support
  • Supports ImageMagick objects directly

17 of 25

Islandora 7 Solution Pack

18 of 25

  • Drupal module installable on top of Newspapers solution pack
  • Newspapers can be segmented
    • On ingest
    • Page-by-page, issue-by-issue
    • Using custom Solr queries

19 of 25

Configuration Page

  • Confidence level is configurable within Drupal
  • Basic authentication is supported via an API key

20 of 25

  • Site owners can run segmentation jobs on subsets of content via Solr queries
  • The Background Processing Drupal module is used to process content without needing to wait on the page

21 of 25

  • The module includes a block to show extracted images next to the main interface

22 of 25

  • Each extracted image is an Islandora object
  • The page contains data like extracted text and confidence score

23 of 25

Islandora 7 to Islandora 2

Why Islandora 7?

  • Our Newspaper archive is in Islandora 7
  • Less of a moving target for new developers c. 2021

24 of 25

Islandora 7 to Islandora 2

To Supporting Islandora 2

  • Extractor API could be used
    • As is with a custom wrapper, or
    • Altered to support Crayfish API
  • Then extracted images are just more derivatives
    • IIIF Manifest can be extended to include sub-sections of a page containing images of various types

25 of 25

Go get it!

All software, including an Islandora 7 VM with sample content, is available on Github:

https://github.com/Islandora-Image-Segmentation