1 of 51

Learning to WALK:

Building a National Web Archiving Collaborative Platform

Ian Milligan

Assistant Professor

@ianmilligan1

Nick Ruest

Digital Assets Librarian

@ruebot

2 of 51

Plan for the Talk

  • Introduction to the Project
  • Workflow
  • WARCLight
  • Derivative Datasets and Providing Access
  • Conclusions

3 of 51

The Current Team

Historian

CS

Librarian/Archivist

The PI/Co-PIs

PhD RA

MA RA

MA RA

Postdoctoral

Fellow (Governance)

4 of 51

The Support

  • Funded by Ontario Ministry of Research and Innovation (all students) and the Social Sciences and Humanities Research Council
  • Computing resources supported by Compute Canada and Microsoft Azure Research Grants program

5 of 51

So what are we all about?

6 of 51

Let’s start with an opportunity not a problem.

7 of 51

We have fantastic web archival collections in Canada.

8 of 51

But not many people use them…

9 of 51

  1. potential users don’t know they exist;

  • current search interfaces don’t aways line up with how scholars work;

10 of 51

(At least until the Archive-It API comes online and transforms everything)

11 of 51

12 of 51

Canadian archival data is silo’d by institution and collection.

13 of 51

The Canadian Web Archival Landscape

  • Currently around 25 Canadian collecting institutions, amassing around ~ 130 collections
  • Most use Archive-It as a back-end provider of web archival services

14 of 51

Right now, to use Canadian web archives – you have to really want to use them.

15 of 51

i.e. you need to be an expert.

16 of 51

We want web archives to be used on page 150 of a random book.

17 of 51

Enter the Web Archives for Longitudinal Knowledge Project

18 of 51

WALKin’

  • Web Archives for Longitudinal Knowledge (WALK)
  • Ian Milligan (Co-PI, UW) + Nick Ruest (Co-PI, York), w/ Geoff Harder, Todd Suomela, Sonya Betz, Peter Binkley, Geoffrey Rockwell (Alberta), Jefferson Bailey (Internet Archive), and John Simpson (Compute Canada).

19 of 51

We want to break down silos, bring Canadian web archives into a centralized portal with access to derivative datasets.

20 of 51

WALKin’

  • Institutions with signed MOUs: Toronto, Alberta, Victoria, Winnipeg, Dalhousie, Simon Fraser University
  • 61 collections (~ half of Canadian web archival collections)
  • 16 TB of WARC files
  • Developing new Solr front end based on Project Blacklight. Data stored on Compute Canada/Dataverse (currently indexed 250 million records).

21 of 51

So what can we do to make sense of all this data?

22 of 51

Plan for the Talk

  • Introduction to the Project
  • Workflow
  • WARCLight
  • Derivative Datasets and Providing Access
  • Conclusions

23 of 51

24 of 51

Workflow

  • Ryan Deschamps (Postdoctoral Fellow and WALK Project Manager)
  • 61 collections, divided in four, divided out to each of the three summer research assistants
  • It’s like jeopardy (“I’ll take HCF Encyclopedia for $100, Dr. Deschamps!”)

25 of 51

In the back end we generate derivative datasets...

26 of 51

Workflow

  • Using Warcbase (http://warcbase.org) and command line tools, we:
    • Transfer data from Internet Archive, run checksums;
    • Generate scholarly derivatives in an automatic, scripted fashion; (visualizations, etc.)
    • Upload some derivatives to Dataverse;
    • Make others available to research team.

27 of 51

So in the front end they…

  • Enter basic information;
  • Describe the collection based on playing with crawl visualizations and hyperlink graphs;
  • Describe the network graph;
  • Use Gephi to lay it out for researchers, and generate pretty pictures;

28 of 51

29 of 51

30 of 51

31 of 51

… and we’ll ultimately.

  • Customizing landing pages, data stored in JSON to generate templated Ruby pages;
  • So that a user can come to the site, read about it, download derivative datasets (hosted in ScholarsPortal Dataverse), and run full-text search.

32 of 51

Landing page

33 of 51

Collection lists

34 of 51

Collection detail

35 of 51

… but … how can they do those searches?

36 of 51

Plan for the Talk

  • Introduction to the Project
  • Workflow
  • WARCLight
  • Derivative Datasets and Providing Access
  • Conclusions

37 of 51

Enter our impending community contribution

  • Last year, we presented on webarchives.ca
  • We’ve indexed ~10T of warcs representing 300M+ Solr docs
    • We grown from one collection!
    • Institution, Collection Name, Collection ID

38 of 51

SCALING

💗 gifcities.org 💗

39 of 51

SolrCloud

40 of 51

Enter our impending community contribution

  • Last year, we presented on webarchives.ca
  • This was a solr index, using the Shine front end
    • Shine is great – but...
    • Doesn’t have an active open source community.
    • We 💖 UKWA

41 of 51

Hello Blacklight

  • Project Blacklight, bringing us into a broader community
  • Better APIs, bug fixes, ability to interact with, give back, communicate with a larger GLAM community (many of you!)

42 of 51

WARCLight

43 of 51

Plan for the Talk

  • Introduction to the Project
  • Workflow
  • WARCLight
  • Derivative Datasets and Providing Access
  • Conclusions

44 of 51

Derivative Datasets

  • For each collection, via Blacklight or just via ScholarsPortal
    • Domain/URL Counts
    • Full text (selected)
    • Network Graphs
  • So a scholar doesn’t need to know about the WARC format, but they can (a) use Gephi; (b) use text analysis; etc.

45 of 51

46 of 51

47 of 51

Suddenly web archives aren’t boutique – they speak to a broader audience and you can imagine how to use them.

48 of 51

Data can then play with other people and platforms

49 of 51

The goal:

One central hub for web archiving search, research, derivatives...

50 of 51

So researchers can cite web archives on page 150 of a book without needing to be an expert!

51 of 51

Thanks!

Ian Milligan

Assistant Professor

@ianmilligan1

Nick Ruest

Digital Assets Librarian

@ruebot