1 of 197

Milan Dojchinovski, Jan Forberg, Johannes Frey, Marvin Hofer, Denis Streitmatter and Sebastian Hellmann and many more

dbpedia.org

DBpedia Knowledge Graph Tech Tutorial

1

2 of 197

Meet the Organizers

Milan Dojchinovski Marvin Hofer Denis Streitmatter Julia Holze

Jan Forberg Johannes Frey Sebastian Hellmann

2

All members of the DBpedia core team hosted by:

Institute of Applied Informatics / DBpedia Association, Leipzig, DE

https://tinyurl.com/DBpediaTechTut

3 of 197

About the tutorial

  • Get familiar with the DBpedia community project
  • Get familiar with the DBpedia Knowledge Graph
    • KG extraction process
    • KG release process
    • KG partitions
  • Learn about the DBpedia Technology stack
    • DBpedia Databus platform
    • DBpedia services: spotlight, lookup
  • Learn how to replicate the DBpedia infrastructure
  • Learn how to make best use of the DBpedia technology

3

https://tinyurl.com/DBpediaTechTut

4 of 197

Agenda

  • PART 1: Getting Started with DBpedia (90 min)
    • Introduction
    • Session 1: DBpedia in a Nutshell (20 min)
    • Session 2: Getting Started (30 min)
    • Session 3: DBpedia`s Blueprint for Creating National Knowledge Graphs (20 min)

Break (15 min)

  • PART 2: Deep into the DBpedia Ecosystem (90 min)
    • Session 4: DBpedia Technology Stack (70 min)
      • Databus, services, release process, id management/fusion, archivo
    • Session 5: Contributions to DBpedia (10 min)
    • Wrap-up / Discussions / Q&A (10 min)

4

https://tinyurl.com/DBpediaTechTut

5 of 197

Guidelines

  • Feel free to ask questions
    • raise a hand, or
    • past a question in the chat, or
    • ask questions during the Q&A slot (after each session)
  • Mute yourself while not speaking
    • … for a better call “hygiene”
  • Turn off your video while not speaking
    • … to save bandwidth
  • The slides will be made public
    • see the footer placeholder: https://tinyurl.com/DBpediaTechTut

5

https://tinyurl.com/DBpediaTechTut

6 of 197

PART 1: Getting Started with DBpedia

6

7 of 197

Session 1: DBpedia in a Nutshell

7

by Milan Dojchinovski

8 of 197

History

2007 - A crowd-sourced community effort to extract structured information from Wikipedia and make this information available on the Web.

  • query Wikipedia as a DB

2021 - Current mission: Global and unified access to knowledge graphs

  • Original definition still holds true, moreover ...
  • Global DBpedia: data beyond Wikipedia
    1. links to recent and authoritative sources
    2. a platform (matchmaking) to integrate your data with all other data
    3. a FAIR {Linked} Data

8

https://tinyurl.com/DBpediaTechTut

9 of 197

The Power of the DBpedia Knowledge Graph

Main SPARQL endpoint: https://dbpedia.org/sparql

SELECT ?person ?name ?country ?population WHERE {

?person a dbo:Person .

?person rdfs:label ?name .

?person dbo:birthPlace ?countryOfBirth .

?countryOfBirth dbo:populationTotal ?population .

FILTER (langMatches( lang(?name), "en" ) )

}

Simple example: “persons, their names in English, their birth country and country population”

9

https://tinyurl.com/DBpediaTechTut

10 of 197

The Power of the DBpedia Knowledge Graph

Main SPARQL endpoint: https://dbpedia.org/sparql

SELECT DISTINCT ?person ?name ?countryOfBirth ?population ?team ?stadium ?stadiumCapacity WHERE {

?person a dbo:Person .

?person rdfs:label ?name .

?person dbo:birthPlace ?countryOfBirth .

?countryOfBirth dbo:populationTotal ?population .

?person dbo:team ?team .

?person dbo:position|dbp:position <http://dbpedia.org/resource/Goalkeeper_(association_football)> .

?team dbo:stadium ?stadium .

?stadium dbo:seatingCapacity ?stadiumCapacity .

FILTER (langMatches( lang(?name), "EN" ) )

FILTER (?stadiumCapacity > 30000)

FILTER (?population > 10000000)

}

ORDER BY DESC(?stadiumCapacity)

More complex example: “soccer players, who are born in a country with more than 10 million inhabitants, who played as goalkeeper for a club that has a stadium with more than 30.000 seats.”

10

https://tinyurl.com/DBpediaTechTut

11 of 197

DBpedia Milestones

11

https://tinyurl.com/DBpediaTechTut

12 of 197

DBpedia is about connecting People and Orgs

12

https://tinyurl.com/DBpediaTechTut

13 of 197

Organizational Structure in Numbers

  • Around 20 DBpedia Chapters
    • language chapters, English, German, Dutch, Czech, Polish, Hungarian, ...
    • regional chapters, e.g. for cities or individual countries
    • domain chapters, e.g. for law, medicine, media and science
    • each chapter hosts and maintains localized DBpedia version
    • more about DBpedia chapters at https://www.dbpedia.org/members/chapter-overview/
  • 32 DBpedia members
    • 41% industry and start-up, 37% non-profit, 22% tiny & self-employed
    • join the network of pioneers to shape the future of knowledge graphs
    • apply via https://www.dbpedia.org/members/membership/

13

https://tinyurl.com/DBpediaTechTut

14 of 197

Organizational Structure

14

https://tinyurl.com/DBpediaTechTut

15 of 197

DBpedia Association Members

Fast growing Knowledge Engineering & Linked Data Lobby

15

https://tinyurl.com/DBpediaTechTut

16 of 197

Overarching DBpedia KG Release Process

  1. Definition of mappings and ontology editing
  2. Execution of the knowledge extraction process over wikipedia dumps
  3. Parsing and validation of the data against strict rules
  4. Release of (intermediate) data artifacts
  5. ID management and knowledge fusion from all language editions
  6. Deployment of the resulting KG

16

1. Mappings, ontology definitions

2. Knowledge extraction

3. Data validation

4. Release of data artifacts

5. ID management and fusion

6. KG Deployment

https://tinyurl.com/DBpediaTechTut

17 of 197

The DBpedia Infrastructure

17

https://tinyurl.com/DBpediaTechTut

18 of 197

Czech DBpedia

18

https://tinyurl.com/DBpediaTechTut

19 of 197

4+2 Main Dataset Groups

Available extractions, 22 billion facts total (500GB without text)

  • Mapping-based (rule-based)
  • Generic (automatic)
  • Text
  • Wikidata

… bonus:

  • Fusion
  • Global IDs

Based on the Wikimedia XML dumps

19

https://tinyurl.com/DBpediaTechTut

20 of 197

Mappings-based Extraction

20

https://tinyurl.com/DBpediaTechTut

21 of 197

Mappings Example

21

{{ PropertyMapping | templateProperty = area_total_km2 | ontologyProperty = areaTotal | unit = squareKilometre }}

{{ PropertyMapping | templateProperty = area_urban_km2 | ontologyProperty = areaUrban | unit = squareKilometre }}

dbr:Prague

dbo:areaTotal 496000000.0 ;

dbo:areaUrban 298000000.0 .

https://tinyurl.com/DBpediaTechTut

22 of 197

Generic Extraction

  • Automatic extraction and export of information
  • Extraction of:
    • unmapped information in infoboxes
    • other structured information found on the Wikipedia pages
  • Automatic extraction of unmapped properties from infoboxes
    • covers all infobox types along with their attributes
    • http://dbpedia.org/property/ + the name of the infobox attribute
    • e.g. http://dbpedia.org/property/birthplace for the Wikipedia attribute “birthplace”
    • objects are created from the attribute values
  • Automatic extraction of other structured information

22

https://tinyurl.com/DBpediaTechTut

23 of 197

Generic Extraction Example

23

dbr:Prague

dbp:name "Prague"@en ;

dbp:nativeName "Praha"@en .

Output triples:

https://tinyurl.com/DBpediaTechTut

24 of 197

Text Extraction

  • Wikipedia articles texts
  • https://databus.dbpedia.org/dbpedia/text/
  • 132 languages, 8 datasets
    • Short and long abstracts
    • content/text + structure
      • sections, sub-sections, paragraphs
      • links
  • Information modeled using the NIF Format
  • Use cases
    • Training data for text mining
    • Fact extraction

Extraction executed every 3-4 months.

24

https://tinyurl.com/DBpediaTechTut

25 of 197

Wikidata Extraction

  • Same approach as for Wikipedia
  • Generic and mappings-based
  • Mappings in JSON

https://databus.dbpedia.org/dbpedia/wikidata

Benefit: Unified access over Wikipedia

and Wikidata

25

https://tinyurl.com/DBpediaTechTut

26 of 197

DBpedia Ontology

  • The heart of DBpedia
  • A shallow cross-domain ontology
    • model information extracted from Wikipedia
    • … BUT goes beyond Wikipedia
    • e.g. mappings for the Dutch National KG
  • Generated on-the-fly
    • when changes in the mappings wiki are introduced
  • Stats
    • over 700 classes and more than 3,000 properties
  • Since v3.7: a directed-acyclic graph, not a tree
    • classes may have multiple superclasses
  • Get it from the Databus

26

https://tinyurl.com/DBpediaTechTut

27 of 197

DBpedia Ontology (cont.)

27

https://tinyurl.com/DBpediaTechTut

28 of 197

DBpedia SPARQL Endpoints

Three core SPARQL endpoints:

  1. DBpedia main SPARQL endpoint
    1. https://dbpedia.org/sparql
    2. hosts the DBpedia latest core release (tiny diamond, see next slide on the KG diamonds)
    3. see https://databus.dbpedia.org/dbpedia/collections/latest-core
  2. Databus SPARQL endpoint
  3. DBpedia Live endpoint

28

https://tinyurl.com/DBpediaTechTut

29 of 197

Innovation: DBpedia KG Diamonds

DBpedia Diamonds: aggregated, ready-to-use, knowledge graphs from Wikipedia/Wikidata and Linked Open Data (LOD).

  1. Tiny diamond (7 million entities)
    • equivalent to DBpedia core/EN hosted in the main endpoint
    • the DBpedia you know since 14 years
  2. Small diamond (140 million entities)
    • a fused graph of 140 Wikipedia languages and Wikidata
  3. Largest diamond (230 million entities)
    • skyrocketing KG, over 220 million entities, 1.45 Billion triples from DBpedia
    • including many LOD sources (DBpedia, Geonames, DNB, Musicbrainz, etc)

For more see: https://www.dbpedia.org/resources/knowledge-graphs/

29

https://tinyurl.com/DBpediaTechTut

30 of 197

DBpedia KG Diamonds Comparison

30

https://tinyurl.com/DBpediaTechTut

31 of 197

Impact of DBpedia

  • In research - over 33,800 articles using or developing technology for DBpedia
  • IBM Watson: DBpedia in the Jeopardy! challenge
  • Web standardisation efforts, e.g. ITS 2.0, SHACL
  • Unicode + DBpedia
    • abbreviations, translations of language codes
  • FactForge.net is a knowledge graph of Linked Open Data (LOD) and news articles about people, organizations and locations.
  • Diffbot’s Knowledge Graph, a large-scale KG extracted from the Web. Uses DBpedia identifiers space.
  • timbr DBpedia: SQL access to DBpedia knowledge: https://www.dbpedia.org/dbpedia-members/timbr/

31

https://tinyurl.com/DBpediaTechTut

32 of 197

Session 2: Getting Started

32

by Jan Forberg

33 of 197

Where can I find DBpedia Data?

33

DBpedia Data

A set of files containing RDF data

https://tinyurl.com/DBpediaTechTut

34 of 197

Where can I find DBpedia Data?

34

OpenLink Virtuoso Triple Store

Option A: The official SPARQL-Endpoint

SPARQL Endpoint

Example:

select distinct ?s where {

?s a dbo:Organisation . �}

YASGUI: http://yasgui.triply.cc/

https://tinyurl.com/DBpediaTechTut

35 of 197

Where can I find DBpedia Data?

35

DBpedia file server

Option B: File Download

Your machine

Download

https://tinyurl.com/DBpediaTechTut

36 of 197

Where can I find DBpedia Data?

36

The DBpedia Databus

DBpedia file server

Option C: The DBpedia Databus

SPARQL Endpoint

Your machine

Download

RDF metadata

https://tinyurl.com/DBpediaTechTut

37 of 197

Where can I find DBpedia Data?

37

The DBpedia Databus

DBpedia file server

Option D: The Latest-Core Collection

Your machine

Download

Databus Collection

https://tinyurl.com/DBpediaTechTut

38 of 197

DBpedia Databus Collections

Use, create and share Databus Collections

What is a Databus Collection?

  • SPARQL-Query + Metadata (Label, Description, etc.)
  • SPARQL-Query selects a set of download URIs using the DBpedia Databus

Why use Databus Collections?

  • Databus Collections are accessible at their own URI
  • Specific data can be selected for applications or tests/benchmarks
  • Collections can be used with the DBpedia Technology Stack
  • No need to write SPARQL
  • Demo: https://databus.dbpedia.org/dbpedia/collections/latest-core

38

https://tinyurl.com/DBpediaTechTut

39 of 197

Where can I find DBpedia Data?

39

Your machine

DBpedia Latest-Core Collection

3 Lines of Bash Script

+

YOU

+

=

Option D: The Latest-Core Collection

https://tinyurl.com/DBpediaTechTut

40 of 197

Where can I find DBpedia Data?

40

Option E: Access through public DBpedia Services

  • DBpedia Spotlight
    • Access at https://demo.dbpedia-spotlight.org/
    • Allows entity annotation of text

https://tinyurl.com/DBpediaTechTut

41 of 197

How to use the DBpedia data?

  • Different tasks require different data
  • Selection of too much data leads to long loading times

Only process the data that you actually need

Option C (Databus) can help

41

https://tinyurl.com/DBpediaTechTut

42 of 197

DBpedia Databus Artifacts

  • Abstract Dataset
  • Points to one or many files
  • Can contain files of different formats, compression types or versions

42

The DBpedia Databus

A File Server

.ttl

.xml

https://tinyurl.com/DBpediaTechTut

43 of 197

Important DBpedia Artifacts

  • Instance Types
  • Labels
  • Geo-Coordinates
  • Mappingbased Objects

43

https://tinyurl.com/DBpediaTechTut

44 of 197

Choosing a Databus Groups

A Databus Group groups multiple Artifacts with common attributes.

DBpedia Artifacts are grouped by Extraction Type:

  • generic: Generic Extraction
  • mappings: Mappingbased Extraction
  • text: Text Extraction
  • wikidata: Wikidata Extraction

44

https://tinyurl.com/DBpediaTechTut

45 of 197

DBpedia Artifact: Labels

Access: https://databus.dbpedia.org/dbpedia/generic/labels/

Tasks:

  • Visualization
  • Label-based indexing

45

https://tinyurl.com/DBpediaTechTut

46 of 197

DBpedia Artifact: Geo-Coordinates

Acccess: https://databus.dbpedia.org/dbpedia/generic/geo-coordinates/

Tasks:

  • Distance-based filtering
  • Map visualizations

46

https://tinyurl.com/DBpediaTechTut

47 of 197

DBpedia Artifact: Instance Types

47

https://tinyurl.com/DBpediaTechTut

48 of 197

DBpedia Artifact: Mappingbased Objects

Access: https://databus.dbpedia.org/dbpedia/mappings/mappingbased-objects/

Tasks:

  • Domain-based filtering
  • Visualisation
  • Data enrichment

48

https://tinyurl.com/DBpediaTechTut

49 of 197

Bonus Artifact: The Databus Itself

Databusception!

49

https://tinyurl.com/DBpediaTechTut

50 of 197

DBpedia Website Demo

50

https://tinyurl.com/DBpediaTechTut

51 of 197

Aggregating multiple Artifacts

51

Your machine

DBpedia Latest-Core Collection

3 Lines of Bash Script

+

YOU

+

=

https://tinyurl.com/DBpediaTechTut

52 of 197

Aggregating multiple Artifacts

52

Your machine

Any Databus

Collection

3 Lines of Bash Script

+

YOU

+

=

The exact data you need for the task at hand

or a comparable amount of LOC in any programming language

https://tinyurl.com/DBpediaTechTut

53 of 197

Session 3:

Dutch National Knowledge Graph

DBpedia`s Blueprint for Creating National Knowledge Graphs

53

by Johannes Frey

54 of 197

National Knowledge Graphs (NKGs)

  • Goal: realize “yellow pages” for Linked Open Data, by integrating national authoritative Linked Data sources in a sustainable, flexible and transparent way
  • serve as global, unified, integrated (high-level), reliable index over the sources to find, discover and study data (FAIR)
  • allow to find entities by combining information of multiple sources, or use as starting point for federated queries
  • study deg. of “dovetailing” of sources
  • “national” scope as a first step
    • authoritative sources with “normative” id space and high/full coverage
    • legal frame drives representation in the ontology / data
    • public bodies with mandate and funding to provide and maintain the data
  • set the foundation for FAIR LOD infrastructure that allows to �integrate sources and create / derive custom KGs [FH21]

54

https://tinyurl.com/DBpediaTechTut

55 of 197

Dutch National Knowledge Graph (DNKG)

  • Pilot 4 authoritative Dutch sources
    • Author thesaurus of the Koninklijke Bibliotheek (KB) the national library of the Netherlands
    • Building and addresses register (BAG) from Netherlands’ Cadastre, Land Registry and Mapping Agency
    • Cultural Heritage Objects (CHO) from the Cultural Heritage Agency
    • Artists Dataset (RKDA) from the Netherlands Institute for Art History
  • 2 common sources / linking hubs with additional information (and links)
    • Dutch DBpedia
    • Digital Bibliography Library Project (DBLP)
  • additional link sources and ID hubs
    • VIAF, Wikidata, Geonames, ORCID

55

DBLP

Dutch DBpedia

Kadaster BAG

Dutch nat. library (KB)

Cultural Heritage Objects

RKD Artists

DNKG

https://tinyurl.com/DBpediaTechTut

56 of 197

DNKG in Numbers & Clustering State

  • Entity Clustering (connected components) based on owl:sameAs links

56

https://tinyurl.com/DBpediaTechTut

57 of 197

DNKG: Analyze Linkage State by Types

  • pie chart: red~unlinked, blue~ext. linked, green~internally linked (based on clusters)
  • arrows portion of int. linked that have DNKG source as target

57

https://tinyurl.com/DBpediaTechTut

58 of 197

Research Questions and Challenges

Creating NKGs:

  • modularization of mappings & transformations for better maintenance and reuse
  • organisation and partitioning of knowledge graphs such that parts can be easily extracted and recombined in a flexible and efficient way
  • transparent, accessible, and tangible workflow that overcomes reliability problems of Linked Data
  • form a central space (on top of decentralized sources) where people can trace down data errors and collaborate

58

https://tinyurl.com/DBpediaTechTut

59 of 197

NKG Creation Blueprint

  • process layer and application layer built on top of DBpedia Databus and Databus Collections
  • intermediate input output artifacts are are managed publicly on the Databus
  • application layer allows interactive / ad-hoc access to results and traceability

59

https://tinyurl.com/DBpediaTechTut

60 of 197

DBpedia Knowledge Cartridges

  • unified, interoperable, modularized, and materialized views of Linked Data sources managed on the DBpedia Databus
  • basic building blocks for a flexible and scalable creation of custom or (National) Knowledge Graphs
  • comparable to blueprinting but for knowledge
  • Deep provenance (original statement values and identifiers)
  • Example: https://databus.dbpedia.org/dnkg/cartridges/rkdartists-cartridge/

60

Pixabay License

https://tinyurl.com/DBpediaTechTut

61 of 197

GFS Data and Provenance Browser (Demo)

  • show all information of one clusters
  • browse between clusters
  • deep provenance (original values, source links)

https://global.dbpedia.org/?s=https%3A%2F%2Fglobal.dbpedia.org%2Fid%2F12Qpnc&p=http%3A%2F%2Fdbpedia.org%2Fontology%2FbirthPlace

61

https://tinyurl.com/DBpediaTechTut

62 of 197

Analytical SPARQL Query on NKG

  • find clusters highly interlinked between DNKG sources

62

https://tinyurl.com/DBpediaTechTut

63 of 197

Greedy Mappings using DBpedia Ontology

  • interoperable KGs by creating mappings to the DBpedia Ontology
  • effort for creating these DNKG mappings seems to follow 80/20 rule

63

https://tinyurl.com/DBpediaTechTut

64 of 197

Direct vs. pt-Construct Mappings

  • direct mappings are managed via DBpedia Mappings wiki http://mappings.dbpedia.org/index.php/OntologyProperty:BirthDate
  • complex mappings/transformations with property-targeted SPARQL Construct Query:
    • target is exactly one DBO property
    • very modular and better reusable
    • efficient execution

64

https://tinyurl.com/DBpediaTechTut

65 of 197

Stay tuned...

Technical details about DNKG follow in as part of the next session “DBpedia Technology Stack”

  • Clustering with Global Identifiers
  • Cartridges
  • Fusion

65

https://tinyurl.com/DBpediaTechTut

66 of 197

Break

15 minutes

66

67 of 197

PART 2: Deep into the DBpedia Ecosystem

67

68 of 197

Session 4:

DBpedia Technology Stack

68

by Jan Forberg, Johannes Frey and Denis Streitmatter

69 of 197

The Databus Technology Stack

The Databus stack is a range of applications and services based on the Databus that allows you to host your own data or build your own applications on top of the stack.

69

Databus Platform

Databus Collections (or SPARQL queries)

on-demand SPARQL stores

on demand

Lookup Search

Mods

Overlay Search

https://tinyurl.com/DBpediaTechTut

70 of 197

The Databus Platform

70

by Jan Forberg

71 of 197

The DBpedia Databus

71

The DBpedia Databus

Some file server

Your machine

Download

Collections

https://tinyurl.com/DBpediaTechTut

72 of 197

Databus - Digital Factory Platform

Registry of files on the Web

  • Global file warehouse
  • Decentralised storage
  • File format doesn’t matter
    • PDF or PDF collection
    • CSV, XML, RDF

72

https://tinyurl.com/DBpediaTechTut

73 of 197

Databus - Digital Factory Platform

… but very strict metadata

  • Provenance (who? - you!)
  • License
  • Granular dataset identity
    • Dataset is a set of files
  • Versioning

73

Rejected Approved

https://tinyurl.com/DBpediaTechTut

74 of 197

Benefits

  • Databus offers DataIds (metadata) and simple file retrieval
  • Mods offer custom extensions of DataIDs
  • Databus Collections offer data selection tools for automated access
  • Ecosystem of DBpedia Docker containers offer easy deployment of a basic Linked Data infrastructure

74

https://tinyurl.com/DBpediaTechTut

75 of 197

The Databus Address Space

Registered data is structured hierarchically (similar to maven repositories). A simple file upload without thought how it could be structured for reuse and maintenance is not possible/allowed. �

75

Publisher (Mike)

Group A (Animals)

Artifact A (Cats)

Artifact B (Dogs)

Version 2

Group B

Version 3

Version 1

Version 3

File F1

type =good

File F2

type =evil

File F1

type =good

File F2

type =evil

File F1

type =great

File F1

type =great

https://tinyurl.com/DBpediaTechTut

76 of 197

The Databus IDs

76

https://tinyurl.com/DBpediaTechTut

77 of 197

SPARQL Access

77

https://tinyurl.com/DBpediaTechTut

78 of 197

Databus - Robot Operations

Load - fully automated in a subscription model

Derive - create a new artifact by conversion, filtering, extraction, textmining, enrichment, statistics, generating a machine learning model, fix errors

Release (Package & Deploy) - packaging means making the data available on the network (e.g. via copying to /var/www Apache2), deploying means uploading the metadata to databus.dbpedia.org

78

https://tinyurl.com/DBpediaTechTut

79 of 197

API Access

TOKEN=$(curl -s -d 'client_id=upload-api' -d 'username=USERNAME -d 'password=PASSWORD -d 'grant_type=password' https://databus.dbpedia.org/auth/realms/databus/protocol/openid-connect/token | cut -d'"' -f 4)

curl -v -X PUT https://databus.dbpedia.org/USERNAME/GROUP/ARTIFACT/VERSION -H "Authorization: Bearer $TOKEN" -d 'DATAID_AS_JSON_LD'

  • Access Token and DataId are validated
  • DataId is added to the Databus
  • BUT: DataId needs to be correct!

79

https://tinyurl.com/DBpediaTechTut

80 of 197

Registering Data via Web GUI

80

81 of 197

Databus Web “Upload”

https://databus.dbpedia.org/system/upload

Required for Databus GUI Upload:

  • Databus Account
  • data (individual files) uploaded on freely accessible web server

81

https://tinyurl.com/DBpediaTechTut

82 of 197

Registering Data with Maven

82

83 of 197

Registering Data with Maven

  • Alternative to Web GUI and API
  • Helps with file organisation
  • Uses decentralized WebId Authentication
  • Authentication mechanism might be subject to change

83

https://tinyurl.com/DBpediaTechTut

84 of 197

Linked WebId and Databus Accounts

WebId and PKCS12 Certificate Tutorial: https://github.com/dbpedia/webid

Databus Account: Register on https://databus.dbpedia.org/

Linking of both accounts: Add your WebId URI to your Databus Account

84

https://tinyurl.com/DBpediaTechTut

85 of 197

Databus Upload with Maven

The cities example project hierarchy (mvn structure)

85

https://tinyurl.com/DBpediaTechTut

86 of 197

Group POM

86

https://tinyurl.com/DBpediaTechTut

87 of 197

Maven Commands

Upload plugin http://dev.dbpedia.org/Databus_Upload_User_Manual with:

  • mvn validate checks account and consistency
  • mvn prepare-package (goal databus:metadata) collects metadata in target/databus/$artifact/$version/dataid.ttl
  • mvn package copies data into a package directory on the server (often /var/www/html/databusrepo/$user/$group/$artifact/$version)
  • mvn deploy posts the dataid.ttl to databus.dbpedia.org

87

https://tinyurl.com/DBpediaTechTut

88 of 197

Databus Collections

88

89 of 197

Databus Collections

89

https://tinyurl.com/DBpediaTechTut

90 of 197

Self-deployable Services

using Docker-Compose

90

91 of 197

Dockerized Databus Applications

Apps are run using docker-compose. Containers usually communicate via volumes.

Download

Container

Application Container

Helper Containers

(e.g. Database)

uses

waits for download

Creates download.lck in target volume

does things once the download is complete

https://tinyurl.com/DBpediaTechTut

92 of 197

(Generic) DBpedia Lookup

https://github.com/dbpedia/dbpedia-lookup

Composite of:

  • Download Container
  • Lookup Container (Application Container)

RDF data is loaded into a graph database to enable SPARQL access�Key-Value pairs are selected via SPARQL queries

92

https://tinyurl.com/DBpediaTechTut

93 of 197

DBpedia Spotlight

93

https://tinyurl.com/DBpediaTechTut

94 of 197

Virtuoso Triple Store

Demo: https://github.com/dbpedia/virtuoso-sparql-endpoint-quickstart

Bash:

git clone https://github.com/dbpedia/virtuoso-sparql-endpoint-quickstart.git�cd virtuoso-sparql-endpoint-quickstart�COLLECTION_URI=https://databus.dbpedia.org/jan/collections/kgc-demo VIRTUOSO_ADMIN_PASSWD=password docker-compose up

Query:

select * where {� ?s ?p ?o .� ?s a <http://xmlns.com/foaf/0.1/Person>�}

94

https://tinyurl.com/DBpediaTechTut

95 of 197

DBpedia Release Process on the DBpedia Databus

95

by Marvin Hofer

96 of 197

The New DBpedia Release Cycle

  • Over 21 B facts per month as RDF triples.
    • Monthly releases of mappings, generic, and wikidata extraction
    • Text extraction every third month
  • MARVIN pre-release and final DBpedia release
  • Improved test methodology
    • JUnit, custom rules, syntax, SHACL, SPARQL tests
    • Minidump and large-scale (final release) evaluation
  • Community Reviewing, due to
    • Process transparency
    • Accessible error reports
    • Direct linking of issues and tests

96

https://tinyurl.com/DBpediaTechTut

97 of 197

The DBpedia Release Cycle

97

START: 10th of month

https://tinyurl.com/DBpediaTechTut

98 of 197

The DBpedia Release Cycle

98

Release raw data

https://tinyurl.com/DBpediaTechTut

99 of 197

The DBpedia Release Cycle

99

Data cleaning

https://tinyurl.com/DBpediaTechTut

100 of 197

The DBpedia Release Cycle

100

Release final data

https://tinyurl.com/DBpediaTechTut

101 of 197

The DBpedia Release Cycle

101

Quality Control

https://tinyurl.com/DBpediaTechTut

102 of 197

Traceability and Issue Management

  • Semantic pinpointing for issue management
  • Explicit association of data artifacts and code (fragments, classes)
  • MiniDump approach for issue management and error handling

102

DataID

Code

File

https://tinyurl.com/DBpediaTechTut

103 of 197

Test driven knowledge extraction

Test Process

  1. Minidump expanded with entities related to data issue(s)
  2. Test definition => SHACL, Construct Validation tests, etc.
  3. DIEF improved and data issues fixed => valid results from testing
  4. Post-release final testing of the introduced improvements

103

https://tinyurl.com/DBpediaTechTut

104 of 197

DBpedia DIEF - Issue Tracker

104

105 of 197

Issue Tracker

105

Please post support questions @forum.dbpedia.org

https://tinyurl.com/DBpediaTechTut

106 of 197

Issue Tracker

106

Please post support questions @forum.dbpedia.org

Decide: If testable by the minidump, � it is data, otherwise software

https://tinyurl.com/DBpediaTechTut

107 of 197

Steps for Data and Software Issues

  1. Create issue based on issue template
  2. Somebody forks the dev branch, write a test, and send pull request (see https://github.com/dbpedia/extraction-framework/.../debug.md)
  3. Check test status (actions API)
    1. Fails? Somebody forks the dev branch, write a fix, and send pull request
    2. Passes? Check test and issue again
  4. Merge into main branch, reference commit hash and close issue

107

Create Issue

Write Test

Write Fix

Close Issue

https://tinyurl.com/DBpediaTechTut

108 of 197

Databus Mods

Metadata Enrichment for Databus Files

108

109 of 197

Databus Mods Overview

109

The DBpedia Databus

Some file server

Databus Mods�master-worker�microservices

Mod Metadata

Federated Queries- find metadata of Databus files�- find data using metadata

metadata access

data flow

https://tinyurl.com/DBpediaTechTut

110 of 197

Databus Mods Metadata

  • realize DataID Extensions (Modifications)
  • allow to link community-specific metadata to the DataID (technical Databus metadata) via Mod Metadata (PROV ontology)
  • Example:
    • FIle is Databus File ID for DBpedia instance types of a specific version
    • Mod Result is additional metadata based on the content of the file, e.g. Online status, MIME-Type, VoID stats
    • Mod Activity Metadata is provenance

110

https://tinyurl.com/DBpediaTechTut

111 of 197

Databus Mods Architecture

  • Microservices with (1) Mod Master, (N) Mod Workers, SPARQL endpoint, and file server

111

Mod Master

  • Polls periodic updates from the DBpedia Databus
  • Schedules configured Mods and dispatches files to workers
  • Collects results from Mod workers
  • Provides results in a SPARQL endpoint and Web Server

Mod Worker (Requirements)

  • Provides an API to be controlled by the Mod Master
  • Implements one specific Mod Metadata generation process�(referred to as Mod Activity)
  • Serves generated Mod Results and Mod Activity Metadata over HTTP

https://tinyurl.com/DBpediaTechTut

112 of 197

VoID Mod Example

  • Vocabulary of Interlinked Datasets
  • RDF dataset insights
    • Total triples
    • Property and Class counts

112

Find RDF datasets containing specific classes or properties

https://tinyurl.com/DBpediaTechTut

113 of 197

DBpedia ID Management,

PreFusion, and Cartridge Creation

113

by Johannes Frey

114 of 197

DBpedia Global Identity Management

114

115 of 197

Global Identity Management

Interoperability Challenges of Linked Data:

  • redundancy in identifiers for same entities and equivalent properties
  • instability of IRI identifiers (changes, link rot, ...)

lead to problems with “qualified references” (I3) and when combining RDF data

→ DBpedia Global IDs (entities) + DBpedia Global Properties for equiv. clusters

https://global.dbpedia.org/id/<base58-ID>

https://global.dbpedia.org/property/<base58-ID>

115

https://tinyurl.com/DBpediaTechTut

116 of 197

Global Identity Management Cycle

  1. harvest links and identifiers from cartridge-inputs
  2. Assigns stable Singleton IDs for external IDs (bijection)
  3. Computes connected components based on owl:sameAs / owl:equivalentProperty
  4. Assigns Global Identifiers based on cluster member

116

https://tinyurl.com/DBpediaTechTut

117 of 197

Global Identity Management Cycle

5. release a new cluster snapshot dump and update microservices which

    • translate external IDs to Singleton & Global IDs
    • allow to discover known references to other datasets for same thing / property
    • NKG: serve fused data as cluster representative → yellow pages for Linked Data

117

https://tinyurl.com/DBpediaTechTut

118 of 197

DBpedia Cartridge Creation

118

119 of 197

Cartridge Creation

1. ID Rewriting

Input from DBpedia Databus:

  • Custom File selection with partially transformed (pt-construct) triples
  • Snapshots of Global ID / Property Assignment (based on Clustering)
  • Replace all IRIs representing local identifiers with their cluster IRI

119

http://fr.dbpedia.org/resource/Paris

4

prop-fr:étages

http://fr.dbpedia.org/resource/Tour_Eiffel

prop-fr:propriétaire

https://tinyurl.com/DBpediaTechTut

120 of 197

Cartridge Combination / PreFusion

2. PreFuse

  • Derive preFused entities by grouping all triples first by same subject and then by their predicate value

120

C1

C2

PreFusion

https://tinyurl.com/DBpediaTechTut

121 of 197

Cartridge / Prefusion Serialization

  • JSON-LD based format tracing down the origin of the triples from all input files (Databus file IDs)
  • compact representation of (alternative) values grouped by subject-predicate pair
  • Unified access to all sources: rewritten global IDs, and properties mapped to property clusters
  • original input statement is recorded via iHash (compressed) or SPO record

121

4 [fr] , 3 [es,en]

dbo:floorCount

https://global.dbpedia.org/id/12HpzV

https://tinyurl.com/DBpediaTechTut

122 of 197

Novel Modular PreFusion

  • Changed from coarse-grained structure above to a partitioning by property clusters (e.g. foaf:name, schema:name) → more flexible selection of “knowledge aspects”
  • 2380 cluster partitions with 90GiB bz2 data and 10 GiB fallback partition for unclustered properties (long tail) for Large Diamond

⇒ useful metadata is a “must” to handle and select these files

122

https://tinyurl.com/DBpediaTechTut

123 of 197

FlexiFusion with Cartridges

  • DBpedia Diamonds: condensed and polished data from multiple sources using Cartridges and FlexiFusion
  • NKGs as incubator for creating a variety of cartridges of the LOD cloud → FAIR LOD vision

123

https://tinyurl.com/DBpediaTechTut

124 of 197

Value sync classification (absolute scale)

124

Blue: value(s) from Source d are synced / not “challenged”

Green: information only in Source d (erroneous / novel??)

Red: all values from Source d are not synced (unique)

Yellow: partially synced but also unique value(s) in Source d � or d is incomplete

  • Sole Source Criterion (SSC): true for an sp-pair of source d if all extracted values contributed from d are only originated in d
  • Alternative Choices Available Criterion (ACC): at least one different extracted value from a source other than d is available in sp

# sp

pairs

sources

https://tinyurl.com/DBpediaTechTut

125 of 197

Evolution: 2019.11 vs 2019.12

125

sources

Blue 0/0

Green 1/0

Red 1/1

Yellow 0/1

Added wikidata -> geonames links

https://tinyurl.com/DBpediaTechTut

126 of 197

Value Sync Classification (normalized scale)

126

https://tinyurl.com/DBpediaTechTut

127 of 197

Mapping Management: Ontology Challenges

  • P1: Which properties/ontologies exist that should be integrated into Mapping Management
  • P2: Need stable, persistent, unified access to ontology versions to create (automatic) mappings between 2 ontologies that can be unsterdood later (provenance, transparency, reproducibility...)

⇒ DBpedia Archivo (Ontology Archive)

127

https://tinyurl.com/DBpediaTechTut

128 of 197

Prefusion Sync Comparison 2019.12 vs 2020.01

128

source

SSC no & ACC no

SSC yes & ACC no

SSC no & ACC yes

Covered source

SSC yes & ACC yes

sum sp

Dnb old

0

101.992.101

0

-

1.356.166

103.348.267

dnb

656.525

97.246.794

434.717

189.061

2.688.940

101.026.976

Musicbrainz old

0

91.852.034

0

-

0

91.852.034

musicbrainz

59

91.748.451

22.523

22.283

72.783

91.843.816

Geonames old

0

160.625.184

362

-

9.745.222

170.370.768

geonames

0

160.445.798

384

326

9.840.852

170.287.034

High values implicate potential incompleteness in

nothing

L,M, compl. source

N ??

N, source errors, property mismatch

But working

all

-

L,M??

L,

Fixed identity management for DNB https IRIs

Musicbrainz identifier dataset IRI fix

Future Vision: incremental Fusion: quantitative Feedback to Linking / Mapping / Normalisation

Blue 0/0

Green 1/0

Red 1/1

Yellow 0/1

https://tinyurl.com/DBpediaTechTut

129 of 197

DBpedia Archivo

129

an Augmented Ontology Archive

by Denis Streitmatter

130 of 197

Problems of Ontologies

Access

  • incorrect RDF deployment
  • link rot / unavailable ontologies
  • different RDF formats
  • no/unclear versioning
  • no stable citation for dependency

Quality

  • missing metadata / documentation (human readability)
  • no/unclear licensing
  • parsing errors (machine readability)
  • logical inconsistencies

130

ID

Accept: application/RDF+XML etc.

e.g. Semantic Web applications

Accept: text/html

e.g. Web browsers

Resource Identifier (NIR)

e.g. http://dbpedia.org/ontology/

https://tinyurl.com/DBpediaTechTut

131 of 197

Archivo Workflow

131

Automatic Ontology Discovery

Ontology Augmentation

Persistence

on the Databus

Ontology

Update

weekly crawl of:

  • ontology repositories
  • classes/properties used on the Databus
  • IRIs used in ontologies
  • user suggestions
  • multiple serialisation formats
  • enhancement with Feature Plugins
  • testing with a SHACL library
  • Star Rating
  • semantic versioning
  • stable abstract identifiers for ontologies
  • (metadata) access via SPARQL/Linked Data
  • crawls every 8 hours

https://tinyurl.com/DBpediaTechTut

132 of 197

Archivo as a Ontology Backup/Citing Tool

  • abstract identifiers for ontologies
  • various parsed serialisations: RDF+XML, Turtle and N-Triples
  • persistent snapshots of any ontology version

One REST request:

132

http://archivo.dbpedia.org/download?

o={ontology-URI}

f={format}

v={version}

e.g. http://archivo.dbpedia.org/download?o=http://datashapes.org/dash&v=2020.07.16-115638&f=ttl

https://tinyurl.com/DBpediaTechTut

133 of 197

Ontology Augmentation & Evaluation

Evaluation:

  • expandable SHACL library for testing ontologies
  • testing the compliance with augmentation features → e.g. proper documentation for LODE
  • testing existence of license → required for Archivo Stars

Augmentation:

  • Multiple parsed serializations of the ontology
  • auto-generated LODE Documentation
  • Semantic versioning - classification of changes (major, minor, patch) based on axiom diff
  • extendable

133

https://tinyurl.com/DBpediaTechTut

134 of 197

Archivo Stars

Baseline:

  • ontology is retrievable without errors
  • some kind of license can be found

Fitness for use stars:

  • license is given with dct:license and is an IRI
  • ontology is logically consistent

134

0x⭐ Ontology

  • not parseable
  • no license provided
  • (maybe) logically inconsistent

2x⭐ Ontology

  • parseable & retrievable
  • some license detected
  • license only human readable or not unified
  • logically inconsistent

4x⭐ Ontology

  • parseable & retrievable
  • unified license URI
  • logically consistent

https://tinyurl.com/DBpediaTechTut

135 of 197

Persistence on the Databus

  • uses the Databus for persistent archiving and access
  • dedicated ontologies agent on the Databus
  • Uses the ontology IRI for identification:
    • publisher → dedicated Databus agent
    • group → domain of the ontology
    • artifact → path of IRI
    • version → timestamp of discovery/update
  • uses multiple metadata properties to annotate ontologies on the Databus
  • persistent, unified versioning & archiving of ontologies
  • access archived metadata via SPARQL / Linked Data

135

https://tinyurl.com/DBpediaTechTut

136 of 197

Access (Live)

  1. Find an ontology → search on the website and on the Databus
  2. Ontology Info Page → visualisation of the previous slides for one ontology
  3. Add an ontology → show an example (http://purl.org/dc/terms/)
  4. Load multiple ontologies in a local SPARQL endpoint

136

git clone https://github.com/dbpedia/virtuoso-sparql-endpoint-quickstart.git

cd virtuoso-sparql-endpoint-quickstart

COLLECTION_URI=https://databus.dbpedia.org/denis/collections/latest_ontologies_as_nt_sample VIRTUOSO_ADMIN_PASSWD=secret docker-compose up

https://tinyurl.com/DBpediaTechTut

137 of 197

Contributions to DBpedia

137

by Milan Dojchinovski

138 of 197

Type of contributions

138

https://tinyurl.com/DBpediaTechTut

139 of 197

Google Summer of Code Projects

This year we will participate in the GSoC project for the 10th time!

Recent GSoC project:

Check more details here: https://www.dbpedia.org/community/gsoc/

139

https://tinyurl.com/DBpediaTechTut

140 of 197

Wrap-up / Q&A

140

by Milan Dojchinovski + all others

141 of 197

What have you learned

  • What is DBpedia
    • knowledge graph
    • knowledge graph technology (platform, processes and services)
    • wide multi-national community of enthusiastic people
    • network of people and organisations
  • How the DBpedia KG is generated and released
    • extraction framework, regular releases
  • How the DBpedia KG is organized
    • data groups, data artifacts
  • Where to find the DBpedia KG
    • databus
  • How to use the DBpedia infrastructure

141

https://tinyurl.com/DBpediaTechTut

142 of 197

Useful pointers

142

https://tinyurl.com/DBpediaTechTut

143 of 197

Join DBpedia

  • Establish DBpedia chapter
  • Become a member
  • Get DBpedia professional services
    • training
    • consulting on your use cases
    • self-hosting DBpedia
    • technical support
    • request material via dbpedia@infai.org

143

https://tinyurl.com/DBpediaTechTut

144 of 197

Thank you! / Q&A

… final thoughts or questions?

144

https://tinyurl.com/DBpediaTechTut

145 of 197

References

  • [ABK07] Auer, S., Bizer, C., Kobilarov, G., Lehmann, J., Cyganiak, R., & Ives, Z. (2007). Dbpedia: A nucleus for a web of open data. In The semantic web (pp. 722-735). Springer, Berlin, Heidelberg.
  • [BLK09] Bizer, C., Lehmann, J., Kobilarov, G., Auer, S., Becker, C., Cyganiak, R., & Hellmann, S. (2009). DBpedia-A crystallization point for the Web of Data. Journal of web semantics, 7(3), 154-165.
  • [FH21] Johannes Frey, Sebastian Hellmann. FAIR Linked Data - Towards a Linked Data Backbone for Users and Machines. WWW '21 Companion
  • [FHO19] - J. Frey, M. Hofer, D. Obraczka, J. Lehmann, and S. Hellmann. DBpedia FlexiFusion the Best of Wikipedia > Wikidata > Your Data, In International Semantic Web Conference, 2019. https://doi.org/10.1007/978-3-030-30796-7_7
  • [FDG20] - J. Frey, D. Streitmatter, F. Götz, S. Hellmann, and N. Arndt. DBpedia Archivo - A Web-Scale Interface for Ontology Archiving under Consumer-oriented Aspects, In Semantic Systems. The Power of AI and Knowledge Graphs, 2020. https://doi.org/10.1007/978-3-030-59833-4_2
  • [FMH19] Johannes Frey, Kay Müller, Sebastian Hellmann, Erhard Rahm, Maria-Esther Vidal: Evaluation of metadata representations in RDF stores. Semantic Web 10(2): 205-229 (2019)
  • [HHD20] Marvin Hofer, Sebastian Hellmann, Milan Dojchinovski, Johannes Frey: The New DBpedia Release Cycle: Increasing Agility and Efficiency in Knowledge Extraction Workflows. SEMANTiCS 2020: 1-18
  • [LIJ15] Lehmann, J., Isele, R., Jakob, M., Jentzsch, A., Kontokostas, D … Hellmann, S. & Bizer, C. (2015). Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia. Semantic web, 6(2), 167-195.
  • [SFM] - S. Hellmann, J. Frey, M. Hofer, M. Dojchinovski, K. Węcel, W. Lewoniewski. Towards a Systematic Approach to Sync Factual Data across Wikipedia, Wikidata and External Data Sources, In Qurator Conference, �Pre-print: https://svn.aksw.org/papers/2020/qurator_gfs/public.pdf

145

https://tinyurl.com/DBpediaTechTut

146 of 197

DBpedia vs. Wikidata

Complementary but still different projects

  • Wikidata not adopted to Wikipedia infoboxes
    • lost of workspace (47k editors vs 13k in Wikidata)
  • Is Wikidata up-to-date?
    • some corona related values we found were/are over 1 year old
    • … likely only for stable values such as birth dates, but not for recent data
  • Wikidata is growing
    • … but this would require a lot more editors to cover all that and keep it updated
    • similar problem with Freebase
  • Live Updates via DBpedia Live
    • whenever something happens, in 30 min in Wikipedia, and then also in DBpedia Live
    • 2 Wikipedia edits every second!
  • DBpedia Global “beyond” Wikipedia
    • link to recent and authoritative sources

146

https://tinyurl.com/DBpediaTechTut

147 of 197

Dockerized Databus Applications

147

148 of 197

Dockerized Databus Applications

Apps are run using docker-compose. Containers usually communicate via volumes.

148

Download

Container

Application Container

Helper Containers

(e.g. Database)

uses

waits for download

Creates download.lck in target volume

does things once the download is complete

https://tinyurl.com/DBpediaTechTut

149 of 197

(Dockerized) SPARQL endpoints

Visit https://github.com/dbpedia/virtuoso-sparql-endpoint-quickstart

Composite of:

  • Download Container
  • Tenforce Virtuoso Container (Helper Container)
  • DBpedia Virtuoso Loader Container (Application Container)

Example: MaStR data loaded into SPARQL endpoint

149

https://tinyurl.com/DBpediaTechTut

150 of 197

DBpedia Lookup

Visit https://github.com/dbpedia/dbpedia-lookup

Composite of:

  • Download Container
  • Lookup Container (Application Container)

Example: https://lookup.dbpedia.org

150

https://tinyurl.com/DBpediaTechTut

151 of 197

Contribute to Archivo

  1. Add ontologies → https://archivo.dbpedia.org/add
  2. Write SHACL tests → https://github.com/dbpedia/archivo/tree/master/shacl-library
  3. Feature Plugins
    1. Suggest them at https://github.com/dbpedia/archivo/issues
    2. Write them yourself with Databus Mods
  4. Find your ontology and fix it → https://archivo.dbpedia.org/list

151

https://tinyurl.com/DBpediaTechTut

152 of 197

DBpedia: Towards FAIR Linked Data

Improvements of Linked Data FAIRness

Databus + Databus Mods

→ F,A,R for data and metadata

Archivo

→ F,A,I,R for ontologies → I for (meta)data

Identity Management

→ I for data and vocabulary properties

152

https://tinyurl.com/DBpediaTechTut

153 of 197

Complex Agents and data workflows on the Databus

153

154 of 197

DBpedia Spotlight

154

https://tinyurl.com/DBpediaTechTut

155 of 197

DBpedia Spotlight

  • Setup DBpedia Spotlight Multilingual
    • Create a volume:
      • docker volume create spotlight-models
    • Create a Docker network:
      • docker network create spotlight-net
    • Download the example docker compose and sites.xml files
      • The docker compose file contains the setup of the DBpedia Spotlight services and the web application.
      • The sites.xml file defines the models available on the web application
    • Run docker compose file
      • docker-compose -f spotlight-compose.yml up -d
    • Spotlight running at: http://localhost:2222/

155

https://tinyurl.com/DBpediaTechTut

156 of 197

DBpedia Spotlight

156

https://tinyurl.com/DBpediaTechTut

157 of 197

Summary

  • DBpedia Databus Platform
    • Upload data on the Databus
    • Work with Databus collections
  • Consuming DBpedia via Databus
    • Dockerized DBpedia
    • Dockerized applications for the DBpedia Stack (Lookup, Spotlight)

157

https://tinyurl.com/DBpediaTechTut

158 of 197

How to use the DBpedia data

Common use cases:

  • Download the data for local processing
  • Create a triple store with the data for SPARQL query access

Possible solution: The docker containers of the DBpedia Technology Stack

158

https://tinyurl.com/DBpediaTechTut

159 of 197

Available Docker Containers

159

TODO Update this…. Option F: Access through self-deployed Apps and Services

  • Same as Option E but with potentially better performance and custom configuration
  • Combines dockerized applications with Databus collections for fast and easy deployment
  • Existing docker containers:
    • Triple Store
    • Triple Store + DBpedia Plugin
    • (DBpedia) Lookup
    • DBpedia Spotlight

https://tinyurl.com/DBpediaTechTut

160 of 197

Data Downloader for Docker Compose

https://hub.docker.com/repository/docker/dbpedia/dbpedia-databus-collection-downloader

Requires:

  • Docker & Docker Compose
  • A DBpedia Databus collection

160

INSERT NICE IMAGE HERE

Container creates a .lock file in the target volume on startup and removes it after the download.

Other containers start to process the data when the .lock file no longer exists.

PRO: No need to code the data access yourself

CON(?): You need to use docker-compose

https://tinyurl.com/DBpediaTechTut

161 of 197

Virtuoso SPARQL Endpoint Quickstart

https://hub.docker.com/r/dbpedia/virtuoso-sparql-endpoint-quickstart

Requires:

  • Docker & Docker Compose
  • A DBpedia Databus collection

161

INSERT NICE IMAGE HERE

Uses the Download-Container and a Tenforce Virtuoso image.

Installer (virtuoso-sparql-endpoint-quickstart) loads the data into the triple store and installs the DBpedia plugin.

Installer and Download-Container terminate and shut down.

https://tinyurl.com/DBpediaTechTut

162 of 197

Data Access on the DBpedia Databus

Demo: Pulling the data with a collection

162

https://tinyurl.com/DBpediaTechTut

163 of 197

Databus - Digital Factory Platform

Build automation tool based on Maven

  • Dataset Identity (ArtifactId)
    • Variance in content/format/compression
  • Optimized for re-releasing the same files
    • 2/3 days to learn and setup the tool (once)
    • 10 minutes to publish an update

https://github.com/dbpedia/databus-maven-plugin

163

Time versioned: 2018.04.10

https://tinyurl.com/DBpediaTechTut

164 of 197

Preliminaries

Manual: User Manual v1.3 · dbpedia/databus-maven-plugin Wiki · GitHub

Tutorial Resources: https://github.com/dbpedia/stack-tutorial-resources

Required for Databus MVN Upload: http://dev.dbpedia.org/Databus_Upload_User_Manual#prerequisites

164

https://tinyurl.com/DBpediaTechTut

165 of 197

Example cities project

165

Publisher (demo)

Group A (stack-tutorial)

Artifact A (cities)

Version

File F1

.ttl

https://tinyurl.com/DBpediaTechTut

166 of 197

Group POM

166

https://tinyurl.com/DBpediaTechTut

167 of 197

Databus Upload

The data project hierarchy:

167

https://tinyurl.com/DBpediaTechTut

168 of 197

Artifact POM

168

https://tinyurl.com/DBpediaTechTut

169 of 197

Databus Upload

The data project hierarchy:

169

https://tinyurl.com/DBpediaTechTut

170 of 197

Artifact Documentation (Markdown)

170

https://tinyurl.com/DBpediaTechTut

171 of 197

Databus Upload

file name prefix needs to match artifact name

folder name needs to be named after the version

171

https://tinyurl.com/DBpediaTechTut

172 of 197

Databus Upload

172

https://tinyurl.com/DBpediaTechTut

173 of 197

Databus Collections

173

174 of 197

Databus Collections

Example collection Open Energy Ontology + MaStR RDF Daten

https://databus.dbpedia.org/jfrey/collections/core-market-data/

Let’s visit: https://databus.dbpedia.org/system/collection-editor

174

https://tinyurl.com/DBpediaTechTut

175 of 197

Databus Access via Website and SPARQL

175

https://tinyurl.com/DBpediaTechTut

176 of 197

Databus Mods & Overlay system

176

177 of 197

Databus Mods

  • realize DataID Extensions (Modifications)
  • allow to link community-specific metadata to the DataID (technical Databus metadata) via Mod Metadata (PROV ontology)
  • Example:
    • Data is Databus File ID for MaStR RDF of wind turbines
    • Mod Result is additional metadata based on the content of the file, e.g. VoiD Mod
    • Mod Metadata is provenance

177

https://tinyurl.com/DBpediaTechTut

178 of 197

VoID Mod

  • “The Vocabulary of Interlinked Datasets (VoID) is concerned with metadata about RDF datasets”
  • can be used to summarize which (ontology) classes and properties are use in an RDF dataset

Example query search for files using the Open energy ontology classes

178

https://tinyurl.com/DBpediaTechTut

179 of 197

Databus Mod Platform

  • Microservice architecture with 1 Master and N Worker services

Mod Server (Master)

  • Polls periodic updates from the DBpedia Databus
  • Schedules configured Mods and dispatches files to workers
  • Collects results from Mod workers
  • Provides results via SPARQL and Linked Data

179

https://tinyurl.com/DBpediaTechTut

180 of 197

Mod Worker

Requirements

  • implements communication API

Benefits

  • decoupling
  • independent of programming language and result format
  • scalable by using, e.g., Docker swarm
  • existing components are easily connectable

180

https://tinyurl.com/DBpediaTechTut

181 of 197

Overlay systems

complex systems that generate consistent metadata overlays can be built on top of Databus + Databus Mods and realize custom services (e.g. an energy asset search).

  • Demo Overlay: Manual Concept Annotation Service
    • provide a list of ontology terms that describe the content of the file
    • https://mods.tools.dbpedia.org:29000/demo/

181

https://tinyurl.com/DBpediaTechTut

182 of 197

Downloading from the Databus

182

183 of 197

Download Options

  • via Databus website
  • Java-based Databus Download client
  • Shell command or R or several other programming languages
  • wget command with http://raw.databus.dbpedia.org
  • Dockerized tools

183

https://tinyurl.com/DBpediaTechTut

184 of 197

Download in R

library(SPARQL) # SPARQL querying package

# Step 1 - Set up preliminaries and define query

# Define the databus endpoint

endpoint <- "https://databus.dbpedia.org/repo/sparql"

# create query statement

q<-"PREFIX dc: <http://purl.org/dc/elements/1.1/>

PREFIX dataid: <http://dataid.dbpedia.org/ns/core#>

PREFIX dct: <http://purl.org/dc/terms/>

PREFIX dcat: <http://www.w3.org/ns/dcat#>

PREFIX xsd: <http://www.w3.org/2001/XMLSchema#>

SELECT DISTINCT ?file WHERE {

?dataset dataid:artifact <https://databus.dbpedia.org/jj-author/mastr/bnetza-mastr> ;

dct:hasVersion '01.04.00'^^xsd:string ;

dcat:distribution ?distribution .

?distribution dataid:formatExtension 'csv'^^xsd:string .

?distribution dcat:downloadURL ?file .

}"

# Step 2 - Use SPARQL package to submit query and save results to a data frame

qd <- SPARQL(endpoint,q)

df <- qd$results

184

https://tinyurl.com/DBpediaTechTut

185 of 197

Raw Databus and wget Download

old-fashioned FTP-like view of the files registered on Databus

https://raw.databus.dbpedia.org/

Download (old) snapshots of Open energy ontology

wget --no-parent --mirror https://raw.databus.dbpedia.org/denis/oe-ontology

185

https://tinyurl.com/DBpediaTechTut

186 of 197

Download with Bash

see dynamic instructions on Databus Collections Website

https://databus.dbpedia.org/jfrey/collections/core-market-data/

186

https://tinyurl.com/DBpediaTechTut

187 of 197

Databus Mod Platform

  • Microservice architecture with 1 Master and N Worker services

Mod Server (Master)

  • Polls periodic updates from the DBpedia Databus
  • Schedules configured Mods and dispatches files to workers
  • Collects results from Mod workers
  • Provides results via SPARQL and Linked Data

187

https://tinyurl.com/DBpediaTechTut

188 of 197

Mod Worker

Requirements

  • implements communication API

Benefits

  • decoupling
  • independent of programming language and result format
  • scalable by using, e.g., Docker swarm
  • existing components are easily connectable

188

https://tinyurl.com/DBpediaTechTut

189 of 197

Overlay systems

complex systems that generate consistent metadata overlays can be built on top of Databus + Databus Mods and realize custom services (e.g. an energy asset search).

  • Demo Overlay: Manual Concept Annotation Service
    • provide a list of ontology terms that describe the content of the file
    • https://mods.tools.dbpedia.org:29000/demo/

189

https://tinyurl.com/DBpediaTechTut

190 of 197

Download Options

  • via Databus website
  • Java-based Databus Download client
  • Shell command or R or several other programming languages
  • wget command with http://raw.databus.dbpedia.org
  • Dockerized tools

190

https://tinyurl.com/DBpediaTechTut

191 of 197

Automatic Ontology Discovery

191

Ontology Repositories

User Suggestions

IRIs of archived Ontologies

VOID summaries from the Databus

Classes/Properties

Ontology IRIs

Ontology Validation

  1. Ont. needs to be typed as owl:Ontology or skos:ConceptScheme
  2. Prevents Ontology Hijacking

https://tinyurl.com/DBpediaTechTut

192 of 197

Augmentation

192

Problem

Archivo Solution

Format Heterogeneity

Deploy as parsed turtle, rdfxml and ntriples

Logical Consistency

Pellet¹ consistency test

Metadata/Human Readability

LODE² documentation, SHACL test

No/unclear License

SHACL tests

(Publishing) Guideline Heterogeneity

Archivo Stars

No/unclear Versioning

timestamps; axiom-based semantic versions

https://tinyurl.com/DBpediaTechTut

193 of 197

Feature Plugins

  • used to augment the ontology with further data
  • can be any service that produces a file/data
  • Possible Examples:
    • Documentation
    • Visualisation
    • Test Reports
  • Implemented:
    • LODE¹ auto generated documentation

193

https://tinyurl.com/DBpediaTechTut

194 of 197

Feature Badges / SHACL Library

  • uses SHACL¹ for testing ontologies
  • Feature Badges → Testing the compliance of Feature Plugins
  • Example: A SHACL test evaluating the compliance to the LODE documentation
    • e.g. a missing label of a class leads to a violation

194

¹https://www.w3.org/TR/shacl/ SHACL → Define the shape of RDF with RDF

Ontology Snapshot on the Databus

Feature Plugin

enhances

checks compliance of ontology with

gets deployed with

gets deployed with

Feature Badge

tested with SHACL

https://tinyurl.com/DBpediaTechTut

195 of 197

Ont. (Re)search and Analysis using Databus Stack

Creating a fresh, local index of the latest ontologies is very easy and automizable

2 easy options:

  1. run docker compose up for Dockerized DBpedia2 using Archivo collection1 and a Virtuoso with all latest ontologies will be deployed automagically
  2. Create a local file dump in a format of your choice with Databus client3

bin/DatabusClient -f nt -s query.sparql (or collection URI)

Shell Guru Tip:

query=$(curl -H "Accept:text/sparql" https://databus.dbpedia.org/jfrey/collections/archivo-latest-ontology-snapshots)

files=$(curl -H "Accept: text/csv" --data-urlencode "query=${query}" https://databus.dbpedia.org/repo/sparql | tail -n+2 | sed 's/"//g')

while IFS= read -r file ; do wget $file; done <<< "$files"

195

https://tinyurl.com/DBpediaTechTut

196 of 197

Flexible Ont. Access via SPARQL (Archivo Metadata)

196

https://tinyurl.com/DBpediaTechTut

197 of 197

Future Work

  • Memento protocol¹ support for Archivo / Databus
  • Utilize more RDF formats for discovery but also as an alternative on the Databus (e.g. the ontology-specific Manchester Syntax)
  • Open contributions for ontologies (reports, validation, extensions (e.g. translation) via Databus Mods Architecture (coming soon)

197

https://tinyurl.com/DBpediaTechTut