1 of 19

Discussion

Speaker

1

FAIR and Open Science in High Energy Physics OAC-2226378, OAC-2226379 and OAC-2226380

2 of 19

Goals

The goals of the workshop will be to:

  1. Assess progress by each experiment in producing reusable data,
  2. Establish updated ideas regarding the use cases for data access, interoperability, and reuse across the different experiments and experimental fields,
  3. Define what data and associated information supports the use cases, and
  4. Identify a preliminary set of access methods and infrastructure that would support these use cases.

2

Technical Recommendations: Provide initial direction for which elements of the cyber ecosystem will be most relevant for the first round of technical improvements. Initiate investigations.

3 of 19

This session

We have 20+45 minutes.

  • Please use Running notes to capture thoughts

3

4 of 19

FAIROS-HEP Scope: Open Science

Yesterday we heard about CERN Open Science Policies. Very broad set of “Domains”

    • Open Access to Publications
    • Open Research Data
    • Open Software
    • Open Hardware
    • Research Integrity, Reuse & Reproducibility
    • Infrastructure for Open Science
    • Research Assessment & Evaluation
    • Education, Training & Outreach
    • Citizen Science
  • Which of the items in the CERN Open Data Policy are in / out scope for FAIROS-HEP?

4

5 of 19

How can FAIROS-HEP Help?

The Implementation Plan for CERN Open Science Policy will be ready soon

  • We mentioned a follow up seminar / meeting to present to community and have follow up discussion.
    • Action Item: Let’s start planning for this. Any roadblocks or complications?
  • Sunje mentioned possibility of a new process to gather community input for future versions of the implementation plan
    • Is this something FAIROS-HEP can help organize?

5

6 of 19

Prompts / Questions

Yesterday we heard repeatedly that effort needed for preservation, open data, reuse, etc. isn’t free and often an unfunded mandate.

  • When we think about the ideal infrastructure for data preservation and reuse, how does this fit into funding of experiments (which largely ends with operations)
  • Should we present the shared infrastructure for preservation and reuse as a facility?

6

7 of 19

Terminology

Yesterday we discussed the DPHEP Level1-4 taxonomy

  • It has been widely adopted, which is a success
  • At the same time, doesn’t naturally describe services that aim to support high-level (Level1) like physics questions but require low-level (Level3 or 4) technical capabilities. Examples:
    • RECAST requires Level4 capabilities of simulation & reconstruction software, but not access to low-level data (only need the Level1 statistical model for background that can be stored on HEPData)
    • LHCb’s ntupleWizard as a service that requires Level3/4 access to data but serves up Level2 to user
    • Cristinel Diaconu will talk to us later today about Open Data vs Preservation
  • It could be useful to develop some terminology / taxonomy to complement the DPHEP taxonomy to help communicate these important nuances.

7

8 of 19

Generating Technical Recommendations

One goal for this meeting is to develop Technical Recommendations:

  • Provide initial direction for which elements of the cyber ecosystem will be most relevant for the first round of technical improvements. Initiate investigations.

Some ideas that surfaced yesterday:

  • Ability to point Binder to HEPData� record so that it prepares an �executable Jupyter environment �(useful for published probability �models, ML models, and �high-dimensional limits stored �as code).
    • You can do this now with Zenodo
    • Or associate Zenodo or GitLab / GitHub �repo with HEPData entry
  • Similar, but for Binder -> REANA

8

9 of 19

Generating Technical Recommendations

One goal for this meeting Technical Recommendations:

  • Provide initial direction for which elements of the cyber ecosystem will be most relevant for the first round of technical improvements. Initiate investigations.

Some ideas that surfaced previously or were generated yesterday

  • Work with INSPIRE to automatically harvest data / code DOIs on zenodo
  • Work with INSPIRE or HEPData to improve, extend, document APIs
  • Work to improve GLANCE (ATLAS) or ALCM (LHCb) -> CAP interface
  • Work to incorporate REANA workflows into ALCM?

9

10 of 19

“Living Publication”

What does it mean?

  • How is it different from / the same as an “executable paper”?

What are the physics use-cases / communities?

  • See Sebastian Neubert’s nice figure �Connecting to Analysis Preservation�and “model-dependence”

What might it look like in 5 years?

  • What do we already have?
  • What is missing?

This is the subject for the afternoon.

10

11 of 19

Other Prompts / Questions

Making HEP Data Fair – what do you need?

Gap Analysis - what are we missing?

Broadening Participation: who is missing?

Ranking Exercise: what should we do first?

11

12 of 19

Pre-workshop Survey Responses

12

13 of 19

Responses from a variety of projects

  • What are the use cases you have seen and foresee for publicly-available data from your organization?
    • Education/outreach, re-use for physics research, re-use for machine learning training, example infrastructure data, inputs for software tutorials
    • Re-use for physics research is happening inside and outside of host collaborations
      • Easier to try new ideas and publish outside of collaboration constraints
  • In practice, what does your organization do to promote re-use of data?
    • Public policies, some framework for sustained funding and development, frequent workshops, regular data releases, substantial and frequently-updated documentation, remind people of funding-agency policies

13

14 of 19

Responses from a variety of projects

  • Please identify missing pieces of infrastructure that would make the generation, dissemination, or reuse of data easier.
    • Additional outreach material, means for federating data repositories, better/more computing resources for users, more widespread data dissemination, better analysis preservation within experiment workflows, better public analysis examples/documentation, more people, sustainable funding

14

15 of 19

Example Thinking around FAIR Principles

15

16 of 19

Findable

The first step in (re)using data is to find them. Metadata and data should be easy to find for both humans and computers. Machine-readable metadata are essential for automatic discovery of datasets and services, so this is an essential component of the FAIRification process.

  • F1. (Meta)data are assigned a globally unique and persistent identifier
  • F2. Data are described with rich metadata (defined by R1 below)
  • F3. Metadata clearly and explicitly include the identifier of the data they describe
  • F4. (Meta)data are registered or indexed in a searchable resource

16

17 of 19

Accessible

Once the user finds the required data, she/he/they need to know how they can be accessed, possibly including authentication and authorisation.

  • A1. (Meta)data are retrievable by their identifier using a standardised communications protocol
    • A1.1 The protocol is open, free, and universally implementable
    • A1.2 The protocol allows for an authentication and authorisation procedure, where necessary
  • A2. Metadata are accessible, even when the data are no longer available

17

18 of 19

Interoperable

The data usually need to be integrated with other data. In addition, the data need to interoperate with applications or workflows for analysis, storage, and processing.

  • I1. (Meta)data use a formal, accessible, shared, and broadly applicable language for knowledge representation.
  • I2. (Meta)data use vocabularies that follow FAIR principles
  • I3. (Meta)data include qualified references to other (meta)data

18

19 of 19

Reusable

The ultimate goal of FAIR is to optimise the reuse of data. To achieve this, metadata and data should be well-described so that they can be replicated and/or combined in different settings.

  • R1. (Meta)data are richly described with a plurality of accurate and relevant attributes
    • R1.1. (Meta)data are released with a clear and accessible data usage license
    • R1.2. (Meta)data are associated with detailed provenance
    • R1.3. (Meta)data meet domain-relevant community standards

19