1 of 63

Open Science and Data Management in Bioinformatics

Caleb Kibet

@Calkibet

2 of 63

@Calkibet

3 of 63

OKF:

@Calkibet

4 of 63

Source: World Wide Web Foundation (2016)

@Calkibet

5 of 63

Africa generates very little research...

@Calkibet

6 of 63

Science is not doing so well

@Calkibet

7 of 63

Science is not doing so well

Lack of replication studies

@Calkibet

8 of 63

Science is not doing so well

Lack of replication studies

Low statistical power

@Calkibet

9 of 63

Science is not doing so well

p-hacking

Lack of replication studies

Low statistical power

@Calkibet

10 of 63

Science is not doing so well

Lack of data Publication bias Paywalls

p-hacking

Lack of replication studies

Low statistical power

@Calkibet

11 of 63

Science is not doing so well

Lack of data Publication bias Paywalls

p-hacking

Lack of replication studies

Low statistical power

HARKing (hypothesizing after results are known)

research excellence rhetoric / games of power and systemic biases in research evaluation and assessment /

lack of trust from the public

@Calkibet

12 of 63

Open Science to the rescue

gather data privately

write journal article

submit for peer-review

peer-review gatekeepers

publish or reject

information available to the public (or not?)

Science 1.0

@Calkibet

13 of 63

Open Science to the rescue

gather data privately

write journal article

submit for peer-review

peer-review gatekeepers

publish or reject

information available to the public (or not?)

pre-register studies

share ideas, methods, protocols, data via blogs, platforms, repositories

submit preprints

publish in blogs, wikis, and in journals

information and data available to the public

Science 1.0

Science 2.0

@Calkibet

14 of 63

So, what is Open Science?

Open Science is the practice of science in such a way that others can collaborate and contribute, where research data, lab notes and other research processes are freely available, under terms that enable reuse, redistribution and reproduction of the research and its underlying data and methods.

(https://www.fosteropenscience.eu/foster-taxonomy/open-science-definition )

@Calkibet

15 of 63

"When all researchers are aware of Open Science, and are trained, supported and guided at all career stages to practice Open Science, the potential is there to fundamentally change the way research is performed and disseminated, fostering a scientific ecosystem in which research gains increased visibility, is shared more efficiently, and is performed with enhanced research integrity."

Open Science Skills Working Group Report (2017)

@Calkibet

16 of 63

@Calkibet

17 of 63

But...the monster of Paywalls and Impact Factors...

@Calkibet

18 of 63

The Paper is the advertisement...

But...it is not always accessible

by Patrick Hochstenbach

CC-BY

@Calkibet

19 of 63

A published article is the tip of the iceberg

“An article (about computational result) is advertising, not scholarship. The actual scholarship is the full software environment, code and data, that produced the result.”

Buckheit and Donoho (1995)

text

data

code

version

science

advertising

reproducibility spectrum

0%

100%

@Calkibet

20 of 63

“When a measure becomes the target, it ceases to be a good measure”

--Goodhart's Law

@Calkibet

21 of 63

"When all researchers are aware of Open Science, and are trained, supported and guided at all career stages to practice Open Science, the potential is there to fundamentally change the way research is performed and disseminated, fostering a scientific ecosystem in which research gains increased visibility, is shared more efficiently, and is performed with enhanced research integrity."

Open Science Skills Working Group Report (2017)

@Calkibet

22 of 63

@Calkibet

23 of 63

Open Science is good for the researcher...

  1. Increased quality of research from reduced errors and fraud due to wider scrutiny and evaluation brought about by transparency
  2. Increased opportunities for both local and global participation in research
  3. Faster transfer of knowledge required to solve problems
  4. Fosters innovation which produces new products and services
  5. Improves productivity and research output due to reduced duplication
  6. Promote awareness among citizens which improves willingness in participation in experiments and data collection.

@Calkibet

24 of 63

Open Science Taxonomy

@Calkibet

25 of 63

@Calkibet

26 of 63

@Calkibet

27 of 63

The monster of Paywalls and Impact Factors...

@Calkibet

28 of 63

The Irony...

“Open Science is transparent and accessible knowledge that is shared and developed through collaborative networks”

@Calkibet

29 of 63

The Paper is the advertisement...

But...

by Patrick Hochstenbach

CC-BY

@Calkibet

30 of 63

@Calkibet

31 of 63

What are your options?

So, even with closed Journals, you can still be Open:

  • Preprint
  • Postprint

@Calkibet

32 of 63

Some Great Options...

Journals with open access and open review:

  • F1000Research
  • The AAS Open Access Journal
  • Wellcome Open Research
  • eLife

@Calkibet

33 of 63

But avoid Predatory Journals...

@Calkibet

34 of 63

Code is the Scholarship...

@Calkibet

35 of 63

Open Science Tools in Bioinformatics

@Calkibet

36 of 63

Git and GitHub

Git is a Version Control System.

  • Helps to keep track of the entire history of things that you are working on.
  • Facilitates collaboration on projects

GitHub is a hosting service for Git Repositories

  • Web-based service for version control and online collaboration
  • The social networking site for developers
  • Used to build a portfolio and get noticed by potential recruiters

@Calkibet

37 of 63

RMarkdown

  • Provides an authoring framework for data science
  • It can be used to save and execute code
  • As well as generate high quality and reproducible reports or presentations that can be shared with an audience
  • It has built-in support for HTML, PDF, MS_Word, RTF, Github, ODT etc

@Calkibet

38 of 63

Literate Programming: Jupyter Notebooks...

  • Open web application for creation and sharing of documents with live code, equations, visualization
  • Notebook documents are human-readable and can contain analysis description, results as well executable code that can be run to perform data analysis.

@Calkibet

39 of 63

Zenodo

@Calkibet

40 of 63

Data repositories

  • Zenodo - is an open access research data repository that provides a place for researchers in any field to deposit datasets up to 50 GB. It has an integration with GitHub to make code hosted on GitHub citable.
  • Figshare - is an online digital repository where researchers can preserve and share their research output i.e. figures, datasets, images and videos.
  • Dryad - is a curated a general-purpose repository that makes the data undisclosed in scientific publications discoverable, freely reusable and citable
  • Dataverse - is an open source web application to share, preserve, cite, explore and analyze research data. Dataverse repository hosts multiple dataverses

@Calkibet

41 of 63

How I’ve practiced Open Science...

@Calkibet

42 of 63

...data, shared to enable and support reuse, would enhance the value of the data generated in the center

@Calkibet

43 of 63

@Calkibet

44 of 63

Why Be FAIR?

  • Save time and increases the value of the data
  • Maximize data discovery and reusability
  • Increase the quality of scientific findings
  • Helps researchers adhere to the expectations and requirements of their funding agencies
  • Frees the researcher to focus on adding value to data rather than on searching, collecting or recreating existing data

@Calkibet

45 of 63

How to be FAIR...

FAIR data requires:

    • a paradigm shift

    • incentive structures and cultural change

    • investments at the Data Management Planning stage

    • Investment in data storage infrastructure and platforms

@Calkibet

46 of 63

FAIR Data ≠ Open Data

"As open as possible, as closed as necessary"

  • Data can be restricted but still be FAIR
  • Open data is just a subset of all the shared data
    • FAIR data should protect the privacy of the data 

@Calkibet

47 of 63

FAIR is controlled data access

@Calkibet

48 of 63

“A goal without a plan is just a wish."� �Antoine de Saint-Exupery

(1900 -1944)

@Calkibet

49 of 63

Data Management Plan

A data management plan describes the data that will be authored and how the data will be managed and made accessible throughout its lifetime."

@Calkibet

50 of 63

What to include in DMP:

  • the types of data to be authored;
  • the standards that would be applied, for example format and metadata content;
  • provisions for archiving and preservation;
  • access policies and provisions

@Calkibet

51 of 63

@Calkibet

52 of 63

Organising your sequencing project

Data Tidiness

  • Think data before you ship out for sequencing - metadata
  • Genomics metadata standards
  • Data in spreadsheets - think computationally:
    • Let alone raw data
    • Row - sample/observations
    • Col - variables
    • Intuitive no-spacing naming
    • Cell - one piece of info
    • Export data in text-format

@Calkibet

53 of 63

Organising your sequencing project

Planning for NGS projects

  • Planning, documenting, organising and execution - reproducibility/replicability
  • Sending out samples for seq - fill in a submission spreadsheet
  • Retrieving results from from facility
    • Documentation - metadata
    • Seq files
    • Validate your data downloads

@Calkibet

54 of 63

Organising your sequencing project

Data storage

    • Store and backup your raw data, it is important that it remains unchanged (file permissions) for fallback - two physically diff locations
    • Minimum - lab head and yourself should have access
    • Storage - HDDs, Severs, Cloud (upto 5G free at OSF)

@Calkibet

55 of 63

Open Science in Bioinformatics

@Calkibet

56 of 63

Bioinformatics is becoming a Data Science, therefore, Open Science Tools should be adopted.

@Calkibet

57 of 63

@Calkibet

58 of 63

Open Science Tools in Bioinformatics

Singularity

@Calkibet

59 of 63

@Calkibet

60 of 63

There is Protocol.io for experimental biologists

@Calkibet

61 of 63

To learn more about Open Science...

@Calkibet

62 of 63

@Calkibet

63 of 63

Our hope is that...

“Future generations [will] look on the term “open science” as a tautology – a throwback from an era before science woke up.

“Open science” will simply become known as science,

and the closed, secretive practices that define our current culture will seem as primitive to them as alchemy is to us.”

-- Brian Nosek and Chris Chambers

@Calkibet