1 of 38

New Models Require

New Action Plans

Implementing Linked Data Within PCC

Steven Folsom, Cornell

<http://orcid.org/0000-0003-3427-5769>

PCC PoCo Meeting

November 1st, 2017

Presentation Link: http://bit.ly/2xJU7Nb

2 of 38

Prompt for Infrastructure White Paper

“...draft a white paper on what the emerging, post-MARC infrastructure might be like… a high-level view of the range of possibilities for that environment. Not a prediction or a prescription of how it ought to be arranged, but an overview of some of the possibilities and the necessary entities and relationships

3 of 38

Acknowledgements

Special thanks to:

  • Linked Data Advisory
  • Chew Chiat Naun (Harvard, $c Mentor and guide)
  • Michelle Durocher (Harvard, $c Mentor and guide)
  • Sarah Ross (Cornell, $c Principal Cataloger and Truth-teller)
  • LD4L (the wolves that raised me in the wilderness)

4 of 38

What is Linked Data?

5 of 38

What from our existing workflows do we want to port forward?

6 of 38

Centralized vs. Distributed Models

7 of 38

Challenges our infrastructure will need to address

8 of 38

discovery (lookups) and caching (copies) of data

9 of 38

Diverse APIs and Data Flows

10 of 38

Diversity of Bibliographic Data Models

11 of 38

Application Profiles and Quality

  • Define the floor, not ceiling
  • “Roomy” enough for all needs
  • wrt RDA, not all data will conform

12 of 38

Linked Data Validation: SHACL as a partial answer

ex:LiteraryWarrantShape a sh:NodeShape ;

sh:scopeClass foaf:Person ;

sh:property [

sh:path foaf:name ;

sh:minCount 1;

sh:maxCount 1;

sh:datatype xsd:string ;

] ;

sh:property [

sh:path ex:createdWork ;

sh:minCount 1;

sh:maxCount 1;

sh:nodeKind sh:IRI ;

] .

:TressieMcMillanCottom a foaf:Person ;

foaf:name “Tressie McMillan Cottom” ;

ex:createdWork :LowerEd .

:Russell a foaf:Person ;

foaf:name “Maria Doria Russell” ;

ex:createdWork “The Sparrow” .

ex:JoshRitter a foaf:Person ;

foaf:firstName “Josh” ;

ex:createdWork :TheAnimalYears .

13 of 38

Reconciliation

<http://orcid.org/0000-0003-3427-5769> ,

owl:sameAs

<http://vivo.cornell.edu/individual/sf433> ,

<http://viaf.org/viaf/316560733> ,

<http://worldcat.org/entity/person/id/2630057950> .

<http://id.loc.gov/authorities/names/no2015079947>

mads:identifiesRWO

<http://orcid.org/0000-0003-3427-5769> .

14 of 38

“Round-tripping”

or... how we’ll need MARC records for the foreseeable future.

15 of 38

Hypothesis: Possible Models and Services

16 of 38

decentralized

publishing

PCC

LD4L

LD4P

IMLS Shared Authorities

OCLC

Ex Libris

Casalini

FOLIO

Hydra

VIVO

Getty

ISNI

id.loc.gov

Orcid

17 of 38

decentralized data discovery as a service

18 of 38

Data flow

PCC

LD4L

LD4P

IMLS Shared Authorities

OCLC

Ex Libris

Casalini

FOLIO

Hydra

VIVO

Getty

ISNI

id.loc.gov

Orcid

19 of 38

What does it mean for

the cataloger?

Not that different

from what we

have now.

20 of 38

Under the hood

we linked to external

resources AND possibly

publish in our own

namespace.

21 of 38

Distributed Vocabulary Maintenance

As a means to be more responsive to current needs and social responsibility.

22 of 38

“Recordless environment”

Manage at the Triplestore level

Named Graphs (like “a record”)

Linked Data Notifications (LDN)

23 of 38

How do we get there?

24 of 38

Safe Assumptions

  • Assume decentralized publication of RDF
  • Assume (at least for the time being) heterogeneous data models
  • Decentralized discovery of RDF will be provided.
  • Support expandable/contractible scope/discovery of the data we link to
  • Understand the many existing APIs for data sources
  • Reconciliation will be needed
  • Profiles and Validation will become important in production environments

25 of 38

The Role of the PCC

26 of 38

Next steps

Evaluate the data models we expect to use, provide feedback

Evaluate the data we would expect to link to

    • Understand the related models so that we can better specify how systems should interact with external data

27 of 38

Next steps

With community and model defined, define shared application profiles.

  • The models aren’t enough.
  • Tune our application profiles through experience.

28 of 38

Next steps

As we consider our desired data models and data sources understand who our partners then become, and what those partnership models should look like.

  • Ideal data models are open, transparent.
  • Ideal data sources are open, transparent.
  • Ideal partnership models are open, transparent.
  • Who’s in our circle of trust and why?

29 of 38

Next steps

Partner with Vendors and Open Source Community to gain experience with early tools (don’t wait for perfection)

    • Define and lobby for functional requirements
    • This includes all metadata platforms, not just linked data platforms

30 of 38

Next steps

Consider how the PCC committee structure and working groups are constructed and communication channels.

31 of 38

The Working Groups

ISNI Pilot

BIBFRAME Task Group

Task Group on URIs in MARC

Task Group on the Work Entity

PCC Steering Committee

Standing Committee on Standards

PCC Policy Committee

Standing Committee on Applications

Standing Committee on

Training

Linked Data Advisory Group

Operations Committees

Task Group on Identity Management in NACO

32 of 38

Organizing the work

33 of 38

Staffing. Who pays for it?

  • Should such important work be staffed almost entirely through volunteerism?

  • How do we train and recruit talent so that we’re equipped to progress conversations the conversations we’re having and solve problems?

  • What is the connection between the work of the PCC and how PCC Member Libraries hire and assign staff?

34 of 38

Training: pet ALL the puppies

35 of 38

Iterative Workplans

36 of 38

Respectfully and with gratitude.

Steven Folsom (Cornell, $c not a guru)

<http://orcid.org/0000-0003-3427-5769>

37 of 38

Image Credits

38 of 38

Image Credits