1 of 93

Building the Next Generation of Web Archive Analysis Service

Panelists: Ian Milligan, Helge Holzmann, Nick Ruest, Samantha Fritz, Thomas Padilla

RESAW 2023

2 of 93

Panel Format

  • RESAW 2023 Conference Website: https://resaw2023.sciencesconf.org/resource/page/id/2
  • Abstract Submission
  • Format: Panel sessions consisting of three individual papers of 15 minutes, introduced by a chair
  • Introduction : Ian
  • Technical Considerations for Building an Arch from Preservation to Research : Helge & Nick
  • Archives Unleashed Cohort Program: Opportunities to Access, Explore, Engage with Web Archives : Sam
  • Beyond Web Archives: Future Directions : Thomas
  • Q + A

3 of 93

Building the Next Generation of Web Archive Analytics Service

Ian Milligan (the panel moderator)

Professor of History & Associate Vice-President, University of Waterloo

4 of 93

Our speakers Today

  • Nick Ruest is an Associate Librarian at York University. He is dedicated to building systems to ensure that valuable historical and cultural materials are preserved and made universally accessible.
  • Helge Holzmann is Web Data Engineer at the Internet Archive. In this role, he directly supports computational research services and is the tech lead for research and development on datasets and analytics for the web archive within Internet Archive and related services in Archive-It.
  • Samantha Fritz is the project manager for the Archives Unleashed Team. She is an information management professional with a passion for open access, information literacy education, and helping people connect with and make sense of data.
  • Thomas Padilla is Deputy Director, Archiving and Data Services at the Internet Archive. He has extensive experience leading collections as data efforts (e.g., Collections as Data: Part to Whole) as well as efforts that advance AI use in cultural heritage organizations (e.g., Responsible Operations).

5 of 93

The common strand is our connection to the Archives Unleashed project.

6 of 93

Archives Unleashed (est. 2017)

The Archives Unleashed project aims to make petabytes of historical internet content accessible to scholars and others interested in researching the recent past. Our team develops web archive search and data analysis tools to enable scholars, librarians and archivists to access, share, and investigate recent history since the early days of the World Wide Web.

Supported primarily by the Andrew W. Mellon Foundation.

7 of 93

Archives Unleashed

Several inter-connected and inter-related projects, which you will learn about today. Some are technical (a few major ones):

  • Archives Unleashed Toolkit (AUT) - 2017 to present
  • Archives Unleashed Cloud - 2017 to 2020
  • Archives Research Compute Hub (ARCH) - 2021 to present
  • Warclight - 2017 to present

Others are events/capacity-building:

  • Datathons (8 events between 2015 and 2020)
  • Cohorts (2 sets of 5 cohorts, 2021-2022 and 2022-2023)

8 of 93

All, however, share the vision of greater access to web archives to enable scholarly exploration at scale.

9 of 93

Our panel today

  • Technical Considerations for Building an Arch from Preservation to Research (Helge Holzmann and Nick Ruest)
  • Archives Unleashed Cohort Program: Opportunities to Access, Explore, Engage with Web Archives (Samantha Fritz)
  • Beyond Web Archives: Future Directions (Thomas Padilla)

10 of 93

Please join me in welcoming our speakers.

11 of 93

Technical Considerations for Building an Archives Research Compute Hub from Preservation to Research

12 of 93

Web archives are all about time travel, right?

13 of 93

Lowering the Barrier to Access:

The Archives Unleashed Project

Nick Ruest (York University)

Ian Milligan (University of Waterloo)

RESAW 2019 - Amsterdam

14 of 93

Archives Unleashed Cloud

  • A web-based front end for working with the Archives Unleashed Toolkit;
  • Runs on our central servers or you can run one yourself;
  • Uses WASAPI – Web Archives Systems API – to transfer data
    • Currently Archive-It supported;
    • We are exploring integration with WebRecorder.io and other WASAPI endpoints
  • Generates a basic set of research derivatives for scholars to work with

15 of 93

What if we made it waaaaaaaay better?

16 of 93

17 of 93

ARCH Platform Details

  • Interactive web application interface both for collection curators and scholarly researchers
  • Generate and download over 20 derivative datasets, and explore them* with Google Colab*
  • UI/UX testing*
  • In-browser visualizations and data previews that presents a glimpse into collection content
  • Located in the Internet Archive data center, ARCH has quick access to the petabytes of content collected

18 of 93

ARCH Stack

18

19 of 93

19

20 of 93

20

21 of 93

21

22 of 93

22

23 of 93

23

24 of 93

24

25 of 93

25

26 of 93

26

27 of 93

27

28 of 93

28

29 of 93

29

30 of 93

30

31 of 93

31

32 of 93

32

33 of 93

33

34 of 93

35 of 93

Multi-tenant job management in ARCH

Helge Holzmann, Web Data Engineer @ Internet Archive

RESAW 2023

36 of 93

Distributed multi-user environment

  • Distributed processing on a Hadoop cluster @ Internet Archive’s infrastructure
    • Job running in parallel on multiple distributed nodes
    • Direct access to Archive-It data and Petabox (archive.org)
  • Powered by IA’s β€œSparkling” data processing library
    • Tuned for efficiency and low memory consumption
  • Advanced job processing and queueing system
    • Multiple queues for different job types
    • Separate queues for sample and full runs
    • Built-in seamless caching layer (LRU)
    • Priority pools prevent jobs from blocking
  • Stream processing to reduce memory

37 of 93

Sparkling ARCH backend

  • IA’s β€œSparkling” data processing library, based on Apache Spark
    • Collection of tools / algorithms to work with (temporal) web data
    • Used internally at IA for various (web) data processing tasks and requests
  • Focus on efficiency and low memory usage
    • Large-scale data processing on distributed computing clusters
    • Designed for efficient web archive data processing
    • Flexible and reusable architecture
      • Development driven by internal and partner requests
      • Supports all common operations we run day to dayοΏ½
  • Open Source: https://github.com/internetarchive/Sparkling

38 of 93

Job management

  • Abstraction layer
    • Hidden away behind the app / APIs
  • Scheduling
    • Multi-user parallel job processing
  • Multiple job types (local / cluster / …)
    • Different queues, multiple each (sample / full)
  • Chained jobs
    • Combining multiple types for post-/processing

39 of 93

Distributed computing

  • Distributed access to the archived data
    • Moving data to the computing nodes
  • Intermediate distributed storage
    • For caching, filtered and special collections
  • Different access strategies
    • Direct WARC access
    • Cached WARC access
    • Random-access through CDX
  • Derivative dataset output
    • Distributed writing

40 of 93

Priority pools

  • Pools of available cluster resources (CPU/RAM)
    • Split among parallel jobs

41 of 93

Priority pools

  • Additional pools after timeout
    • New jobs can bypass long-running jobs
    • Long-running jobs keep running at lower priority

42 of 93

Priority pools

  • After high-priority jobs have finished
    • Low-priority jobs can claim back all resources

43 of 93

Data processing

  • Filtering down input datasets
    • Reducing load, increasing efficiency
  • Sub-collection access via fast CDX pointers
    • Touching only required records
  • Advanced job-dependent sampling
    • Generating representative example outputs
  • Stream processing
    • Every record is read only once as it gets streamed
    • Minimizing memory consumption

44 of 93

Archives Unleashed

Cohort Program

Opportunity to Access, Explore,and Engage with Web Archives

Samantha Fritz

RESAW 2023, 5-6 June 2023, Marseille, France

Back

Next

45 of 93

Anticipating opportunities to expand ARCH’s scope to facilitate responsible use of all digital collection content types

Tool Development

Overview of technical specs for building ARCH, an novel, deeply-integrated self-service compute hub for processing web archive collections

Research Engagement

Reflections on Archives Unleashed Cohorts, a mentorship program aimed at fostering research engagement with web archives

Future Developments

Back

Next

Nick Ruest + Helge Holzmann

Samantha Fritz

Thomas Padilla

46 of 93

Archives Unleashed Project

To make petabytes of historical internet content accessible to scholars and others interested in researching the recent past, by focusing on lowering the barriers to access and use of web archives, specifically by building user friendly analytical tools, accessible web archival interfaces, and community infrastructure and engagement.

Back

Next

Archives Unleashed

(2017 -2020)

How it started

Merge Archives Unleashed with the Internet Archive Archive-It Platform to create an end-to-end solution to collect and study web archives.

Foster and support a research community of practice by offering opportunities to engage with web archive research.

Archives Unleashed + Internet Archive

(2020 -2023)

How it’s going

47 of 93

Archives Unleashed Project

Back

Next

Archives Unleashed + Internet Archive

(2020 -2023)

How it’s going

Archives Unleashed

(2017 -2020)

How it started

Outcomes

Outcomes

Cohort Program

Archives Unleashed Toolkit (1.2.0)

Archives Research Compute Hub

Notebooks

Archives Unleashed Toolkit

Archives Unleashed Cloud

Datathon Series

48 of 93

The Cohort Program

Back

Next

50 Researchers

20 Institutions

9 Countries

49 of 93

The Cohort Program

Back

Next

  • Year long intensive collaborations
  • Supports research with web archives as a primary data source and scholarly objects
  • Technical support
  • Mentorship from Archives Unleashed team, field expertise, and peer-to-peer support
  • Pilot users of ARCH β†’ generate datasets on collections for more thorough analysis

50 of 93

Web Archives + Research

01

Back

Next

51 of 93

The Web

The Web

Back

Next

  • has been archived since mid-1990s and led to an abundance of material
  • provides a rich context for research data (scope & scale) 1
  • remains an untapped resource for research 2

1. Milligan I. History in the Age of Abundance?β€―: How the Web Is Transforming Historical Research. McGill-Queen’s University Press; 2019.

2. Schroeder, R., & BrΓΌgger, N. (2017). Introduction: The web as history. In R. Schroeder & N. BrΓΌgger (Eds.), The Web as History: Using Web Archives to Understand the Past and the Present (pp. 1–20). UCL Press. https://doi.org/10.2307/j.ctt1mtz55k.6

Image Description:

hairball network analysis visualization produced through Gephi

52 of 93

Back

Next

Acute challenges faced by scholars:

  • available analytics tools
  • community infrastructure
  • inaccessible / technical burdensome interfaces
  • time & resources
  • and…….

404 - File Not Found

53 of 93

Web Archive Data is...complicated

Back

Next

While web archives provide a rich source of primary data…..

WARC data is messy, complex, and difficult to work with

54 of 93

Access

02

Back

Next

55 of 93

Access

Back

Next

  • Without access, there is no use
  • Bridge gap between researchers and data
  • Cohorts given ARCH account for analysis
  • Direct, read-access to Archive-It collections
  • Building relationships with data owners & curators

56 of 93

Explore

03

Back

Next

57 of 93

Explore

Back

Next

Image Credit: John Tenniel, Public domain, via Wikimedia Commons

  • Embracing challenges, and the unfamiliar, complex nature of web archives
  • Majority of researchers, new to web archives and limited experience in large-scale computational methods
  • Empower curiosity and experimentation
  • Skills-building, knowledge sharing through research consultation and peer-to-peer support
  • Teams experimented with tools and approaches that appeal to broader research skill sets and experience levels

58 of 93

Explore

Back

Next

  • Embracing challenges, and the unfamiliar, complex nature of web archives
  • Majority of researchers, new to web archives and limited experience in large-scale computational methods
  • Empower curiosity and experimentation
  • Skills-building, knowledge sharing through research consultation and peer-to-peer support
  • Teams experimented with tools and approaches that appeal to broader research skill sets and experience levels

59 of 93

Engage

04

Back

Next

60 of 93

Engage

Back

Next

  • Build relationships between collection curators and researchers
    • 23 Archive-It Partners
    • 44 web archive collections
  • Contribute to visibility of institutional collections
  • Highlight the possibilities of how web archives can be investigated

61 of 93

Crisi Communication in the Niagara Region during the COVID-19 Pandemic

Using the COVID-19 in Niagara collection, this team examined how organizations in the Niagara region communicated about bylaw changes, masking requirements, and the vaccine rollout in the first two years of the COVID-19 pandemic.

Methods: Close reading, text analysis, attention to temporal shifts

Tools/Resources: SolrWayBack, Jupyter Notebooks

Additional: Testing curriculum and pedagogical methods to support computational use of web archives by undergraduate students

Back

Next

Image Credit: Tim Ribaric, et al. via Crisis Communication in Niagara, β€œOur Google Colab notebooks”

Tim Ribaric, David Sharron, Cal Murgu, Karen Louise Smith, and Duncan Koerber

62 of 93

AWAC2 Analysing Web Archives of the COVID Crisis through the IIPC Novel Coronavirus dataset

Back

Next

This team investigated transnational events to better understand actors, content types, interconnectivity, and representativeness found within the IIPC Novel Coronavirus web archive collection.

Methods: Distant reading, text analysis, topic modelling,

Tools/Resources: Pandas, Iramuteq, Latent Dirichlet Allocation (LDA), Word2vec and Doc2vec

ValΓ©rie Schafer, Karin De Wild, FrΓ©dΓ©ric Clavert, Niels BrΓΌgger, Susan Aasman, Sophie Gebeil, Joshgun Sirajzade

63 of 93

Latin American Women’s Rights Movements: Tracing Online Presence through Language, Time and Space

Back

Next

Using web archive collections related to human rights and feminist movements, the project will develop a historical analysis of the websites from women’s rights movements in Mexico and Latin America, particularly those focusing on eradicating femicides and gender violence.

Methods: Topic Modelling (w/ geo, linguistic and temporal contrast) Image analysis,

Tools/Resources: Jupyter Notebooks

Additional: collection, training, knowledge-sharing, and advocacy of activist movements and digital memory of Spanish-speaking voices are preserved

Image Credit: Fernandez, S., et al. Map of web archiving initiatives in Latin America

Sylvia Fernandez, Rosario Rogel-Salazar, VerΓ³nica BenΓ­tez-PΓ©rez, Alan ColΓ­n-Arce, Abraham GarcΓ­a-Monroy, Hejin Shin

64 of 93

Cohort Projects

  • Crisis Communication in the Niagara Region during the COVID-19 Pandemic by Tim Ribaric, David Sharron, Cal Murgu, Karen Louise Smith, Duncan Koerber
  • Everything Old is New Again: A Comparative Analysis of Feminist Media Tactics between the 2nd- to 4th Waves by Shana MacDonald, Aynur Kadir, Brianna Wiens, Sid Heeg
  • Mapping and tracking the development of online commenting systems on news websites between 1996–2021 by Anne Helmond, Johannes Paßmann, Robert Jansma, Luca Hammer, Lisa Gerzen
  • AWAC2 β€” Analysing Web Archives of the COVID Crisis through the IIPC Novel Coronavirus dataset by ValΓ©rie Schafer, Karin De Wild, FrΓ©dΓ©ric Clavert, Niels BrΓΌgger, Susan Aasman, Sophie Gebeil, Joshgun Sirajzade
  • Viral health misinformation from Geocities to COVID-19 by Shawn Walker, Michael Simeone, Kristy Roschke, Anna Muldoon
  • Latin American Women’s Rights Movements: Tracing Online Presence through Language, Time and Space by Sylvia Fernandez, Rosario Rogel-Salazar, VerΓ³nica BenΓ­tez-PΓ©rez, Alan ColΓ­n-Arce, Abraham GarcΓ­a-Monroy, Hejin Shin
  • Querying Queer Web Archives by Filipa Calado, Corey Clawson, Di Yoong, Lisa rhody
  • Web Archiving and the Saskatchewan COVID Archive: Expanding Coverage to Capture Social Media, Medical Misinformation, and Radicalization by Jim Clifford, Derek Cameron, Erika Dyck, Craig Harkema, Patrick Chasse, Tim Hutchinson, Ahmad Rahman
  • Historicizing Aughts-Era Mormon Mommy Blogging Media Landscapes by Emily Edwards, Robin Hershkowitz, Lauren Andrikanich
  • Using Web Archives for Mapping the Use of Cultural Practices in Postconflict Societies and During Reconciliation Processes by Ricardo Velasco Trujillo and Luis Gomez

Back

Next

65 of 93

Takeaways

Back

Next

Connect with tools; embrace challenges and the unfamiliar

Connect with data; forge relationships with data curators/creators

Access

Explore

Connect with communities to share experiences and knowledge

Engage

66 of 93

Back

Next

Thank You!

Archives Unleashed Project: https://archivesunleashed.org/

Twitter: @UnleashArchives

Documentation

Archives Unleashed Toolkit

Archives Research Compute Hub (ARCH) User Documentation

Archives Unleashed Cohorts Reading List

Credits: This presentation template was created by Slidesgo, and includes icons by Flaticon, and infographics & images by Freepick

67 of 93

πŸ€–πŸŽ¨πŸ”ŠπŸ“šπŸŽžοΈπŸ€–

Thomas Padilla

Deputy Director, ADS

Internet Archive

ARCH

Future Directions

68 of 93

69 of 93

The Now

The Future

70 of 93

The Now

The Future

71 of 93

With ARCH, researchers and educators can make use of vast amounts of data contained within web archives to gain insights into a wide range of subjects, spanning historic and contemporary events. ARCH’s intuitive interface makes it easy to build research-ready datasets and dataset publishing capabilities ensure that research findings can be shared with others in a transparent and reproducible way.

72 of 93

Research and Education Service

73 of 93

~ 8 years of shared experience between AU & IA supporting researchers in this space

20 institutions, 50 researchers, 9 countries using ARCH for research through Mellon supported AU & IA collaboration

~25 GLAM institutions piloting how they’ll use ARCH to support researchers and … in some cases collection curation

Development directly informed by

74 of 93

75 of 93

76 of 93

custom multi-institutional collection building

for analysis

77 of 93

publish and preserve

78 of 93

79 of 93

πŸš€

June 26, public service launch!

80 of 93

The Now

The Future

81 of 93

82 of 93

GLAM

collections as data

83 of 93

GLAM

collections as data

84 of 93

researchers & GLAM practitioners

85 of 93

ARCH will expand to support ingest, derivation, and analysis of any digital collection content type (e.g., image, text, audio, etc.)

πŸŽ¨πŸ”ŠπŸ“šπŸŽžοΈ

86 of 93

Content scope to include IA and non-IA collections

87 of 93

Develop ability for ARCH to streamline work for small, medium, and large organizations. Preliminary set of partners identified - more would be great - please reach out!

88 of 93

Expansion of job types to support diversified set of data job needs according to content type and research domain (e.g., object recognition in artworks, color composition, etc.)

πŸ§‘β€πŸŽ¨πŸŽ¨

89 of 93

Machine learning and AI integrations

90 of 93

Move to host notebook server rather than rely on Google Colab (we know where that road leads πŸ™ƒ)

91 of 93

Partner on cultivation and support of expanded set of user communities focused on particular areas of work that require computational methods at scale

92 of 93

πŸ‘πŸ‘πŸ‘

Samantha Fritz, Ian Milligan, Nick Ruest

Derek Enos, Alex Dempsey, Helge Holzmann, Jefferson Bailey, Kody Willis, Karl Blumenthal, Bridget Collings, Sylvie Rollason-Cass, Tegan Broderick, Catherine Falls, Lori Donovan

93 of 93

πŸ€–πŸŽ¨πŸ”ŠπŸ“šπŸŽžοΈπŸ€–

Interested in ARCH?

Have an idea for partnership?

Please reach out!

thomas padilla

tpadilla@archive.org