1 of 16

Educating Information Science Professionals on Digital Language Archiving

Presentation at the Computational Resource for South Asian Languages (CoRSAL) VIII Symposium, October 4, 2024

Dr. Oksana L. Zavalina, University of North Texas

2 of 16

Presentation Outline

  1. Digital language archives’ user needs
  2. Filling the gap in training for librarians and archivists in stewardng language archive collections: developing the specialized graduate course
  3. Preliminary results and next steps, community engagement
  4. CoRSAL VIII Symposium participants feedback?

2

3 of 16

General Needs of Digital Language Archive Users

General information needs do not depend on the specific use purpose:

  • The main general need is discoverability of information resources (achieved through providing quality metadata)
  • Metadata should support find, identify, select, obtain, and explore user tasks.

Some examples of digital language archives users’ general needs are for services available:

  • Stream-able audio and video with transcriptions and translations;
  • User interface accessible on mobile devices;
  • Downloadable machine-readable text files and bulk download option.

Burke, Zavalina, Chelliah, & Phillips (2022)

3

4 of 16

Specific Needs of Digital Language Archive Users

Community users have specific needs for supporting language revitalization (Burke, 2023 findings with Boro language community):

  • Resources available:
    • Dictionaries
    • Textbooks and teaching aids for different subjects
    • Storybooks for children; Folktales, stories
    • Resources on religion and culture
    • Attractive items that capture attention
  • Services available:
    • Grouping resources by reading level, grade level, etc.

​

4

5 of 16

Current State of Meeting These User Needs

Challenges for linguists interacting with digital language archives:

  • Observations (Burke et al., 2022) identified deficiencies in completeness and accuracy of metadata as an important challenge:

5

    • “This is a text...it's just a scan, not really accessible in any way other than hand transcription. So, in the cases where text does exist, it would be useful to denote the type of text. Some PDFs are searchable, and some are not.”
    • “There are so many things tagged as text that are just an audio file....if I needed data, how would I systematically gather data? it would be a little difficult for me, given just the disparity of formats and availability.”

6 of 16

Current State of Meeting These User Needs (1)

Language community members have reported access barriers (focus group):

​

Academics are building these archives...and so you build it for people like yourself. So, the door is an academic door, right? So other academics walk along and say, “Oh! I know how to open this door. And it’s for me! And everything in there is for me!”

And for other people who are not academics, they look at these archives and they’re like looking at tools from some foreign thing...the door isn’t made for them

(Wasson et al., 2016, p. 675)

6

7 of 16

Preparing Information Professionals to Help Support Language Archive User Needs

Laura Bush 21st Century Librarian Program grant-funded project RE-254860-OLS-23

  • collaboration between CoRSAL digital language archive, UNT College of Information, UNT Libraries, IU Linguistics Department, with contributions from Chin Languages Research Project
  • Developing (2023-2025) open source modular curriculum with a strong practical component to educate information professionals in the archiving and curation of resources that provide the means to revitalize community memory and language.
  • Learning materials suitable for use in:
      • courses offered in Library & Information Science, Archival Studies university programs
      • on-the-job trainings of library/museum/archive staff working with language collections
      • archiving workshops for members of indigenous, immigrant & refugee communities who are documenting their community heritage

7

8 of 16

Interdisciplinary Project Team

  • Co-PI Dr. Shobhana Chelliah (Linguistics, Indiana University)
  • Co-PI Dr. Mark E. Phillips (Library & Information Science: digital repositories, University of North Texas - UNT)
  • PI Dr. Oksana Zavalina (Library & Information Science: information organization, UNT)
  • Dr. Ana Roeschley (Archival Studies, UNT)
  • Dr. Brian C. O’Connor (Library & Information Science: digital Imaging, UNT)
  • Graduate students (Information Science + Linguistics):
  • Sergio I. Coronado, Merrion Frederick, Hugh Paterson III at UNT
  • Alexandra O’Neil at IU

8

9 of 16

Community Language Archiving and Curation for Information Professionals Course: Overall Structure

9

Learning modules

Learning objectives (top-level)

M1: Planning, developing, and managing a community language archive

  1. Define community language archives and their functions.  
  2. Describe procedures and considerations for community language archive planning and development.  

M2:  Ethical archival practices and digital curation in community language archives.

  1. Examine archival theory and practice and digital curation trends and perspectives relevant for community language archives.  
  2. Explain important access, preservation, and description issues and practical problems associated with community language archives.  

M3:  Metadata, digital content management and web archiving for community language archives

  1. Identify metadata standards that can be utilized in representing community language archive materials to support information needs and user tasks of community language archive users.  
  2. Describe the problems related to selection of digital content management tools, depositing and web archiving for community language archives.  

M4: Dissemination and use of community language archive content, evaluation of archival services

  1. Identify efficient and ethical ways of disseminating the content of community language archives.  
  2. Discuss approaches for the evaluation of the services provided by community language archives.  

10 of 16

Project-Developed Course Testing at UNT

Target audience: current and future UNT graduate students in Master of Library Science, Master of Information Science, & Ph.D. in Information Science programs (including but not limited to Archival Studies concentration, Linguistics concentration)

10

Developing learning materials, getting ADA & copyright compliance approvals

1st offering:

​

Summer 2024 5-week semester as a section of INFO 5680 seminar, with 11 federally-funded stipends (23 students)

2nd offering:

​

Summer 2025 10-week semester as INFO 5385 specialized course, with 14 federally-funded stipends

​

11 of 16

Training with Strong Experiential Component �(all 4 assignments are practical or real-life analytical): One Example

In Module 1, students plan, create -- by following best practices for language archives -- and document the process of creation of a mini-collection of 4 born-digital & digitized archive items based on student’s own 5-words wordlist:

  • 1 image, 1 text, 1 video, & 1 audio

11

Summer 2024 students’ collections represented 12 languages/dialects and various topics:

12 of 16

Student Feedback on the Project-Developed Course is Very Positive: Some Examples

12

“I learned so much about both community archives and community language archives. I was unaware that there was a community for either of these and now it’s a big, new world that I would love to be involved with.”

“I really did enjoy getting hands-on experience and being able to interact”

​

“My concentration is [...] but after this course I'm wishing I had chosen Archives!”

​

“It was informative and interesting. I wish I could have done this in a 16 week semester”

13 of 16

Some Findings of Course Surveys for Summer 2024

  • The level of learning-objective-related confidence developed by students:
    • Meets & exceeds our expectations for course-level objectives (91%-100%)
    • Overall, is as expected for module-level objectives (70%+)
  • Learning objectives that were addressed not only by instructor presentations but also by module assignments, tended to be met by higher percentage of students

13

14 of 16

Next Steps Include Obtaining Community & Professional Feedback and Improving the Project-Developed Learning Materials Based on it

Project Advisory Board (experts in documentary linguistics, digital language archive creators, etc.):

  • Dr. Kristine Hildebrandt (Endangered Language Fund)
  • Dr. Andrea Berez-Kroeker (Kaipuleohone Digital Language Archivewaii)
  • Dr. Susan Kung (Archive of the Indigenous Languages of Latin America, developer of the Archiving for the Future training)
  • Dr. Jeonghyun Kim (director of the Digital Curation program at UNT)
  • Cristela Garcia-Spitz (Melanesian and Pacific islander studies, UC San Diego Libraries)

14

Speakers and learners of endangered languages:

  • Dr. Kelly Berkson (IU, Chin Languages Research Project)
  • Dr. Kenneth Van Bik (Hakha Lai Language, California State University, Fullerton)
  • Members of Hakha Lai & other language communities

15 of 16

Next Steps: Continued

  • Develop 10-week & 16-week versions of the course with more detailed content and more tasks in practical assignments
  • Refine project-developed learning content for integrating in other courses for information professionals, including:
    • UNT INFO 5224 Advanced Metadata (Modules 2-3 continued integration since Spring 2024)
    • UNT INFO 5960 Cultural Heritage Stewardship (Module 2 continued integration since Spring 2024)
    • UNT INFO 5742 Web Archiving (Module 3 integration planned for Spring 2025)
  • Finalize project-developed learning materials and make available for anyone interested as open source CC BY NC 4.0 materials (through UNT Scholarly Works repository)

​

15

16 of 16

Thank you!

​

What do you think?

Please share your feedback with our project team:

In-person during the CoRSAL VIII (2024) Symposium

and/or any time via contact information on the project website

16