1 of 25

EarthCube FAIR How-to Series:

Data Identifiers: Anchoring your data in times of change

Nancy J. Hoebelheinrich

EarthCube Office & Knowledge Motifs LLC

November 2021

This work is supported through the NSF award #1928208.

2 of 25

What is a key challenge to conducting research in the current data ecosystem?

  • Constant flux in requirements from:
    • Funders
    • Publishers
    • Repositories
    • Sponsoring organizations

Image by Bruno /Germany from Pixabay

3 of 25

How can you more easily meet the challenges of changing requirements with one [fairly] straightforward strategy?

Image by David Mark from Pixabay

4 of 25

The strategy: Assigning IDs to your data

Here you will learn:

  • How data IDs can help you reduce your research workflow
  • How data IDs can improve the broader research ecosystem
  • Questions to ask to figure out which ID type to use for your data

Tools you will find:

  • An ID Assignment Checklist for your dataset
  • Examples of different types of datasets & scenarios where the Checklist is used to navigate the acquisition of IDs for your data

5 of 25

What if…

  • Dataset Type Example 1:
    • Static dataset (no longer being updated, but still associated with published research results)
      • e.g., CSV file, table or figure representing a view of or from the dataset

You are ready to publish your findings and need to associate your static dataset (or subsets) with your paper…

- What do you need to do?

- How do you do it?

6 of 25

What if…

  • Dataset Type Example 2:
    • “Dynamic” dataset (actively updated, but available for citing in order to facilitate analysis -- the results of which may then be published),
      • e.g., time series data

A colleague on a collaborative research project wants to publish findings from the project and needs an ID to associate a snapshot of the continually updated dataset...

- What do you need to do?

- How do you do it?

7 of 25

What if…

  • Dataset Type Example 3:
    • Field collection dataset (used as the basis for subsequent and comparative analyses)
      • e.g., Physical soil samples from which physical, chemical & biological measures will be derived

You’re involved in a cross- disciplinary environmental research project for which you want to integrate data from multi-modal analyses derived from a physical sample...

- What do you need to do?

- How do you do it?

8 of 25

What’s a common denominator for all these dataset types & use scenarios?

  • All will require assigned Identifiers -- but, for what purposes?

  • Some key purposes for identifiers:
    • Describe unique entities
    • Point to persistent data location
    • Associate relationships among members of a collection (including parent/child, location, or event)

9 of 25

Identifiers – What are they?

    • Unique identification codes applied to digital and physical “objects” that need to be unambiguously referenced
    • Identifiers can be assigned to people, organizations, publications, data of various types, models, code, standards, instruments, assays, and more
    • Some identifiers are or should be persistent Identifiers
      • Persistent identifiers (PIDs) are essentially “permanently” assigned and maintained by a trusted organization, making the IDs “permanently” actionable over a reasonable time period.

10 of 25

How do IDs for data help you?

    • Provide stability and consistency in a changing data storage & usage ecosystem
    • Give you accurate, current & consistent links to the location of your (data) and their metadata even if they move to different storage environments
    • Associate your data with publications, related authors, relevant organizations, other research objects such as entities in databases, information systems, and knowledge graphs (for credit and reuse!)

11 of 25

How do IDs help in the broader research data ecosystem?

    • Facilitate citation of your data so that others can find, retrieve and refer to them within the scholarly communication and publishing arena
    • Provide links to related research products that facilitate data integration
    • Foster reproducibility of your research via identification & location of parent/subset versions
    • Are key to enabling longer term access, preservation and archiving of your data

12 of 25

To get started…

Identify colleagues who can help you:

  • Fellow researchers in your subject domain
  • Research data specialists at your institution
  • Support staff at the organization where you plan to deposit your data

Confirm that you (as the researcher) can register IDs for your dataset by:

  • Checking to see if IDs already exist
  • If IDs don’t already exist, checking with the organization that is helping you deposit your data to make sure that you can register the IDs

13 of 25

To get started…we’d like to help

We’ve prepared a detailed ID Assignment Checklist that you can use to gather information about yourself, your research project and your data that you will need to choose and acquire a data identifier

14 of 25

ID Assignment Checklist: Which ID scheme should I use?

  • Ask the research support staff at your institution what are the standard ID schemes used by your community or data domain.
  • Check to make sure that the ID schemes used by the community are flexible enough to use for the types of datasets you and your collaborators produce.

15 of 25

ID Assignment Checklist: Which ID schema should I use?

  • Check data similar to yours to see whether the ID schemes have been successfully applied to the use cases that you anticipate for your data. For example, can the IDs be used :
    • At varying levels of granularity (per Dataset Type 2)?
    • Or used to reflect relationships among disparate data entities (per Dataset Type 3)?

16 of 25

ID Assignment Checklist: Which ID should I use?

  • Ask if the infrastructure and services of the allocation agent that provides and registers the IDs has sustained organizational & community support. For ex.,
    • Do the data repositories or archives where you plan to store your dataset for the longer term support these IDs?
    • Do the publishers that you intend to use accept these IDs?
    • Does the registration agent provide training or help with the description (metadata) requirements or other technical questions?

17 of 25

ID Assignment Checklist: Preparing for the ID assignment

  • Pull together the important information required by the ID scheme to describe your data & to make your data more FAIR (Findable… Accessible… Interoperable…Reusable)
    • What’s important for finding your dataset in a catalog or archive and making it accessible?
    • Include the contextual information most important for others to reuse your dataset & to be more easily interoperable

18 of 25

Once you’ve chosen an ID scheme for your dataset, keep in mind...

  • You have a role to play in making the IDs work over time as does the data archive or repository you’ve chosen
  • For an identifier to work as it is intended, especially a persistent ID, all parties involved have responsibilities and are making a commitment…

19 of 25

The commitment

  • Should be supported by a “landing page” that can give the larger context for the dataset by including, e.g., :
    • An actionable link to download the dataset
    • A preferred citation including the identifier
    • Version number
    • Listing/linking to other versions, parent/subsets or other related data

  • IDs must come with important information (metadata) to be effective
    • Descriptions of and about your dataset at the range & depth needed to be useful -- which may vary depending on the chosen ID and its fitness for use / purpose
    • The dataset location to which the chosen ID points

20 of 25

Whose commitment?

Those who are in the best position to know should submit & keep the metadata and the landing pages up to date:

    • Primarily the data repository where your data are stored
    • But could be you, the researcher

21 of 25

See ID Assignment Checklist Dataset Type Examples

For Dataset types

Learn about:

  • Best IDs for your type of data
  • Where to register your dataset with an ID & what kind of descriptive info you’ll need
  • Repositories that accept your chosen ID

22 of 25

In Summary: One strategy to help you anchor your research data & workflow in times of change

We’ve discussed:

  • How assigning data IDs can help you reduce your research workflow
  • How data IDs can improve the broader research ecosystem
  • Questions to ask to figure out which ID type to use for your data

You’ve seen:

Examples of different dataset types & how an ID Assignment Checklist can be used to get or assign IDs to your dataset

23 of 25

References and Resources

24 of 25

Acknowledgements

For those reviewing, commenting and making suggestions on the content of this presentation, I’d like to thank:

  • Kristin Vanderbilt, Information Manager at Florida Coastal Everglades LTER Program, Florida International University
  • Melissa Cragin, Data Strategist, EarthCube Office, San Diego Supercomputer Center, University of California, San Diego
  • Megan Carter, Community Director, Earth Science Information Partners (ESIP)

Copyright: Nancy J. Hoebelheinrich

25 of 25