1 of 41

If you are looking for the presentation at the Thematiendaagse Metadata (May 21, 2021),

please go here: https://tinyurl.com/KIA20210521

2 of 41

Project SKILLNET

Het delen van kennis in het heden én verleden

Kennisnetwerk Informatie en Archief

Kennisplatforms E-depot en Toegang

webinarreeks “Toegankelijkheid organiseren”

Ingeborg van Vugt en Liliana Melgar

15 December 2020

This presentation (15 December 2020) is available here:

https://tinyurl.com/KIA-SKILLNET

3 of 41

Vandaag….

Deel 1. Beschrijving SKILLNET project

  • Onderzoek naar de Republiek der Letteren

Deel 2. Verzamelen metadata brieven

  • Brieven metadata
  • Crowdsourcing: CEMROL
  • Catalogus Epistularum Neerlandicarum

Deel 3. Challenges metadata (Catalogus Epistularum Neerlandicarum)

  • Obtaining the data
  • Standardizing dates
  • Harmonizing person names

Deel 4. Challenges full text (ePistolarium)

Conclusie en toekomst

3

4 of 41

Vandaag….

Deel 1. Beschrijving SKILLNET project

  • Onderzoek naar de Republiek der Letteren

Deel 2. Verzamelen metadata brieven

  • Brieven metadata
  • Crowdsourcing: CEMROL
  • Catalogus Epistularum Neerlandicarum

Deel 3. Challenges metadata (Catalogus Epistularum Neerlandicarum)

  • Obtaining the data
  • Scoping the corpus
  • Standardizing dates
  • Harmonizing person names

Deel 4. Challenges full text (ePistolarium)

Conclusie en toekomst

4

5 of 41

SKILLNET

  • European Research Council (Dirk van Miert, 2018-2022)
  • Ideaal van kennisdeling in de vroegmoderne gemeenschap van geleerden in Europa = Republiek der Letteren (1500-1800)

Sharing Knowledge in Learned and Literary Networks

Universiteit van Utrecht, geesteswetenschappen

https://skillnet.nl/

6 of 41

Wat is de Republiek der Letteren?

Erasmus, oil on panel by Hans Holbein the Younger, 1523–24; in het Louvre, Parijs https://www.britannica.com/biography/Erasmus-Dutch-humanist

Brief van Francesco Malocchi aan Carolus Clusius, UBL, VUL 101 (1606) © University Libraries Leiden

Tekening van de Narcissus in een brief van Clusius aan Matteo Caccini, 1608-10-10 © University Libraries Leiden

Brief van Pietro Guerrini aan Cosimo III (01-07-1691) , Il viaggio in Europa di Pietro Guerrini (1682-1686). Edizione della corrispondenza e dei disegni di un inviato di Cosimo III dei Medici

Brief van Antonio Magliabechi aan Gijsbert Cuper, 02-04-1685, KB, MS 72 D 10

Magliabechi aan Jacob Gronovius, ongedateerd, LMU 778, f. 78

7 of 41

SKILLNET

  • Vijf onderzoeksprojecten

o Van Miert: Republiek der Letteren in de moderne tijd

o Llano: Prosopografie en sociale structuur

o Scholten: Herinnering en identiteit van geleerden

o Hollewand: Concept Republiek der Letteren

o van Vugt: Netwerken Republiek der Letteren

  • data specialist (Liliana Melgar)
  • project assistant (Robin Buning)

Brieven Staatsarchief Florence, Antonio Magliabechi

8 of 41

Vandaag….

Deel 1. Beschrijving SKILLNET project

  • Onderzoek naar de Republiek der Letteren

Deel 2. Verzamelen metadata brieven

  • Brieven metadata
  • Crowdsourcing: CEMROL
  • Catalogus Epistularum Neerlandicarum

Deel 3. Challenges metadata (Catalogus Epistularum Neerlandicarum)

  • Obtaining the data
  • Scoping the corpus
  • Standardizing dates
  • Harmonizing person names

Deel 4. Challenges full text (ePistolarium)

Conclusie en toekomst

8

9 of 41

9

Plaats en datum van verzending

ontvanger

Plaats van ontvangst

afzender

Inhoud

Metadata

10 of 41

10

Netwerk Analyse (persoonsnamen, datum)

Bron: Catalogus Epistularum Neerlandicarum

11 of 41

Bronnen

Metadata en full-text

11

Correspondentie Antonio Magliabechi (1633-1714) �circa 22.000 brieven

12 of 41

CEMROL

Collecting Epistolary Metadata from the Republic of Letters

https://cemrol.hum.uu.nl

  • Gedigitaliseerde gedrukte (vroegmoderne) brieven edities
  • Open Domein beschikbaar
  • Vrijwilligers helpen ons de metadata van vroegmoderne brieven te verzamelen (crowdsourcing)
  • Publieksdagen

12

13 of 41

13

Belangstelling? https://cemrol.hum.uu.nl

14 of 41

14

Stap 1: Markeren

15 of 41

15

Stap 2: Transcriberen (persoonsnamen)

16 of 41

16

Stap 2: Transcriberen (datum/plaats)

17 of 41

17

Interface:

https://tinyurl.com/CEMROL-interface

18 of 41

18

EPISTOLARIUM

CEN

Datasets, via verschillende platforms wereldwijd

19 of 41

De Catalogus Epistularum Neerlandicarum (CEN)

“De CEN is een catalogus van de tot op heden ontsloten brieven berustend in de collecties van samenwerkende Nederlandse erfgoedinstellingen. De CEN bevat ongeveer 585.000 beschrijvingen van losse brieven en correspondentiedelen geschreven in de vroegmoderne en moderne tijd (ca. 1500 - heden). Daarmee verschaft de CEN informatie over in totaal meer dan 2 miljoen brieven. Het bestand wordt regelmatig uitgebreid.”

19

  • Ontstaan in c. 1984/5 - data toegevoegd tot 2019?
  • Brieven van 1303 tot 2003 (XML dump KB Oktober 2019: 584.723 records)
  • Hoeveel brieven? Aantal brieven rows (Collection-level/ item-level)
  • Onderhouden bij de KB, later door OCLC Leiden, nu in Worldcat.

20 of 41

Catalogus Epistularum Neerlandicarum (CEN)

(access via Picarta tot 2020, via WorldCat vanaf 2021)

20

21 of 41

De Catalogus Epistularum Neerlandicarum (CEN)

eerste stap: een grote schoonmaak.

  • Waarom CEN?
    • Gericht op collecties, niet op individuen (meta-archief, data-silos)
    • Nooit eerder bestudeerd als een corpus

  • Uitdaging:
    • Het is “messy
      • Veel verschillende bibliothecarissen/archivisten hebben data toegevoegd over vele jaren.
      • Bedoeld om collecties te vinden (curation-driven), niet als onderzoeksdataset (research-driven)

21

22 of 41

Vandaag….

Deel 1. Beschrijving SKILLNET project

  • Onderzoek naar de Republiek der Letteren

Deel 2. Verzamelen metadata brieven

  • Brieven metadata
  • Crowdsourcing: CEMROL
  • Catalogus Epistularum Neerlandicarum

Deel 3. Challenges metadata (Catalogus Epistularum Neerlandicarum)

  • Obtaining the data
  • Scoping the corpus
  • Standardizing dates
  • Harmonizing person names

Deel 4. Challenges full text (ePistolarium)

Conclusie en toekomst

22

23 of 41

Obtaining the metadata

  • In general, it is not easy for researchers to find “catalogs” as “research datasets” in which the metadata is ready to be used
    • Catalog metadata is created for “findability” (retrieval) purposes
    • Catalog metadata is created by many persons over time
    • The cataloguer cannot do extensive research to find out missing data
    • Letter standards have not always existed and are not consistently applied:
  • Obtaining the CEN catalog:
    • Ingeborg got the data via OCLC (library organization), permission given by the Nationale Bibliotheek (KB)
      • Row XML file, complete, no filters
      • Post-processed to convert it to CSV
    • Difficulties:
      • ISNI numbers (identifiers for persons) not included in original export
      • Column with company names missing in original export (but these names were still present in the ISBD description)

23

24 of 41

Preparing the metadata

Data preparation:

  • Scoping it to the research questions
    • Tailoring the corpus to correspond to the period of interest
  • Making the data consistent, accurate, and reliable and for answering research questions
    • The key metadata fields are present and prepared (document type (letter or correspondence), letter dates, titles, persons (sender, receiver), places)

24

Data preparation => Research quality

  • Preparing data has to be a transparent process
  • It has to facilitate accountability: researchers have to be able to account for the transformations that have been done to the data)

Image by Yannis Velegrakis (Utrecht Data Science webinar, October, 2020)

25 of 41

Tools used in the CEN data preparation process

  • Excel => not used, quite limited for the task

  • OpenRefine => intensively used, very powerful, lots of functionality:

    • Faceting and filtering
    • Querying (also using regular expressions)
    • Replacing (also using regular expressions)
    • Joining, splitting
    • Integrating small scripts (e.g., to create unique identifiers, or to replace using REGEX, to map different datasets and extract data)
    • Common/fast transformations
    • Clustering! A wide range of advanced algorithms available at the click of a button
    • Exporting all or parts of the data, reimporting, merging, combining…
    • And, most importantly: it keeps track of the steps! → Accountability

25

https://openrefine.org/

“A free, open source, powerful tool for working with messy data”

26 of 41

Tools used in cleaning CEN

26

  • Python
    • Pandas library (to work with “data frames” and powerful data structures)
    • Libraries specialized on string matching (FuzzyWuzzy, JellyFish)

  • Jupyter notebooks
    • to write the code/scripts and the descriptions of what is being done
    • as a way to provide narrative to the code, and making the scripts easier to be shared

  • Google’s Collaboratory
    • to share the Jupyter notebooks with researchers in the cloud (no installation of Python or Anaconda needed)

27 of 41

Scoping the corpus

  • Defining criteria
    • start date 1303 (no filter needed)
    • final date: 1815 (end of the Napoleonic period)
  • Strategy:
    • Parse year string to decade
    • Leave out letters with dates after 1820 (5 year buffer for uncertain dates) but include everything from earliest year
    • Leave in all rows with missing date (to check later if they were part of the period of interest)
    • letters from/to persons that were born after 1850 are excluded
  • Resulting initial number of rows -> 66.067
    • But, careful! rows ≄ letters

27

Initial parsing of “Jaar” field per decade to make a selection of the sub-corpus

Sentinel for missing value

28 of 41

Parsing the data: e.g., cleaning the “aantal” column

Data preparation very important step: properly parsing (each column has a single type of value)

28

Original: 568 choices,

mixed data types

Cleaned--> 4 columns

  • We can see that there are around 50 thousand individual descriptions for one letter (thus, c. 50.000 letters).
  • But there are c. 15.000 descriptions of “brieven”, thus we may never know what the final amount of letters in the CEN catalogus is!

Plus

  • A column for “Type of document” (extracted from title column)
  • Extra descriptions (e.g., “gedigitaliseerd”, “doos”, etc., are not separated from original column → not relevant for research questions

Original: Type of document is part of the title

29 of 41

Cleansing the data: e.g., date fields

Cleansing letter and persons’ dates

  • Creating homogeneous data types
    • e.g., in a date field, a year such as “1582?” or “ca.1582” is not desirable, but this mark should be moved to an extra column (e.g., “uncertain date”)
  • Clustering variants (syntactic clustering)
    • e.g., “28 mei 1582”, and “1582, mei 28” become one
  • Clustering variants (semantic clustering)
    • For example, the dates above may not be the same if we consider the two different calendars:
      • 18 mei 1582 (Gregorian)
      • 28 mei 1582 (Julian)

29

Original letter date field: 15383 values

Big date ranges

Unknown dates

Uncertain dates

Known month/day

but unknown year

30 of 41

Cleansing the data: e.g., date fields (2)

  • 9545 (14% of the total descriptions) have specific dates

30

Cleansed version: 2 columns (range)

Plus extra information columns:

  • JaarIsRange
  • InferredDateCEN
  • InferredDateSkillnet
  • JaarCleanedSkillnet

31 of 41

Working with strings: Harmonizing Persons’ names

  • Ideal situation
  • Be able to know which letters were written by one person

  • But real initial data from CEN:
  • One column for Sender, one column for Receiver
  • Different forms of the name in each of those columns

Common issues in working with person names:

  • Synonyms
    • For example, in the Epistolarium project, they found that the name "Christiaan Huygens" was spelled in ca. 300 different ways
  • Homonyms
    • For example, we have at least three different Christiaan Huygens in CEN:
      • Christiaan Huygens (1551-1624)
      • Christiaan Huygens (1629-1695)
      • Christiaan Huygens (1634-1676)
  • Lack of unique identifiers
    • CEN PersonID inconsistently assigned
    • not every person/entity has an ID)
    • there are some cases in which a different ID was assigned to the same person)
    • there are no other identifiers for persons in the current dataset

31

32 of 41

Clustering name variants

32

Or, one could also merge the dates and CEN identifiers before doing the clustering:

Harmonization strategy

  • Combine all names into an “authority list” of person names (with roles for “afzender/ontvanger”)
  • Parse dates of birth/death and other information from the person name (string)
  • Data imputation to persons’ dates (e.g., “floriat” date based on letter date)
  • Cluster variants using the combination of:
    • String matching script (for name string)
    • Mapping rules based on dates
  • Remove duplicates (making sure to store them as variants)
  • Reconciling person names with external sources
    • (e.g., Wikidata, VIAF, CERL thesaurus, EMLO persons, Biographisch portaal, etc.)
  • Provide unique identifiers as much as possible, hopefully estable/shared IDs

Tools

  • OpenRefine has powerful “clustering” algorithms, but cluster variants based on strings only:

But, it is more powerful and efficient to use a Python mapping script that uses a combination of string matching scores and rules based on dates

33 of 41

Vandaag….

Deel 1. Beschrijving SKILLNET project

  • Onderzoek naar de Republiek der Letteren

Deel 2. Verzamelen metadata brieven

  • Brieven metadata
  • Crowdsourcing: CEMROL
  • Catalogus Epistularum Neerlandicarum

Deel 3. Challenges metadata (Catalogus Epistularum Neerlandicarum)

  • Obtaining the data
  • Scoping the corpus
  • Standardizing dates
  • Harmonizing person names

Deel 4. Challenges full text (ePistolarium)

Conclusie en toekomst

33

34 of 41

The case of ePistolarium

  • Circulation of Knowledge and Learned Practices in the 17th-century Dutch Republic
  • In The Netherlands, Epistolarium shows how can letters be searched, browsed and (to a certain extent) be used for research
  • The CKCC corpus consists of various correspondences of scholars who were active in the Netherlands in the 17th-century. It currently consists of approximately 20,000 letters.”

34

http://ckcc.huygens.knaw.nl/epistolarium/

CSV file download

(with metadata only)

One of the main targets of this project is to create free, online access to historical sources, open to researchers from various disciplines all over the world. “

35 of 41

Obtaining the full text (content)

35

  • However, access to the letters’ full text is not possible for the entire corpus, not beyond individual letters (one at the time)
  • There are very valuable experiments with named entity recognition, but.. they cannot be used for corpus-driven research (e.g., doing content analysis of all the letters by Hugo de Groot).

Or yes, you can, but then you have to

  • Ask the archive directly (they provide a data dump that is made available via “Nederlab”, more suitable for linguistic analyses), or
  • ...scrape the website (not ideal!)

36 of 41

Working with text: analyses based on concepts

  • SKILLNET researchers who are interested in the use of the term “republic of letters” have to deal with more than 190 permutations of this term in Latin (see blog post: https://skillnet.nl/the-intricacies-of-conceptual-history-the-192-republics-of-letters-in-latin/)
  • CTRL-F is used often, to have the sense of more control and detailed inspection
  • Or, more flexible regular expressions can also be used (see screenshot below)
  • But, online environments, designed for searching and browsing, don’t allow for the flexibility needed when doing basic (not to say advanced) text analysis using tailored scripts. Downloading the content is then needed for research, but because of common restrictions, providing more flexible ways (tools) to work with data are needed (e.g., Jupyter notebooks that use the provider’s APIs in a protected environment --see the CLARIAH Media Suite, for example).
  • At SKILLNET we use Jupyter notebooks internally, for doing research and learning:

36

Regular expression for the variations of the word “Republic”

37 of 41

Vandaag….

Deel 1. Beschrijving SKILLNET project

  • Onderzoek naar de Republiek der Letteren

Deel 2. Verzamelen metadata brieven

  • Brieven metadata
  • Crowdsourcing: CEMROL
  • Catalogus Epistularum Neerlandicarum

Deel 3. Challenges metadata (Catalogus Epistularum Neerlandicarum)

  • Obtaining the data
  • Scoping the corpus
  • Standardizing dates
  • Harmonizing person names

Deel 4. Challenges full text (ePistolarium)

Conclusie en toekomst

37

38 of 41

DATA uitdagingen voor onderzoekers (samenvatting)

  • Metadata en full-text uit verschillende soorten bronnen
  • Data zijn verspreid over verschillende platforms, wereldwijd
  • Het uitwisselen van data is lastig voor onderzoekers, data is niet altijd toegankelijk
  • Grote variatie persoonsnamen (geen IDs), plaatsnamen en datering in brieven.
  • Er zijn een paar standaarden voor het catalogiseren van brieven metadata (Correspondence Metadata Interchange format, EMLO templates), maar vaak geen gedetailleerde aanbevelingen.
  • Geen common identifiers voor brieven (e.g. DOI, ISBN boeken)

38

39 of 41

Lessons learned during data preparation

  • Once more, the importance of “traceable/non destructive cleaning” (always duplicate columns that you will clean and leave the original intact)
  • Try to use identifiers for persons (and letters or other documents) every time it is possible
    • Look for Person identifiers for example in Wikidata, VIAF, WorldCat Identifies, CERL, EMLO
  • Researchers should try to use templates for describing Persons and Letters (or other documents)
  • Cataloguers could contribute to initiatives in improving metadata standards for letters
  • Research data curators should document, as much as possible, the changes applied to the data, and the process itself (to remember what you did, and to report it back when you publish or present)

39

40 of 41

Future steps: sharing and making data available for others to do research

40

  • We aim to finalize cleaning CEN and making it ready as research dataset for SKILLNET researchers in early 2021:
    • Harmonizing and reconciling place names
    • Mapping and finding overlaps with other letter datasets (e.g., CEMROL, EMLO)

  • At the end of the project, in 2022, we aim to make CEN, and other research data, available for consultation and reuse:
    • Via Nodegoat web platform (SKILLNET instance of Nodegoat, only for project members until the end of the project)
    • Depositing it in a repository after the project ends
      • Repository options:
        • Dans (national repository for research data)
        • Yoda (institutional Utrecht University repository)
        • e-Depot (state archives repository)

41 of 41

References

41