1 of 10

Vocabularies in TERN

Junrong Yu

​

​

Software Engineer in TERN

junrong.yu@uq.edu.au

2 of 10

Terrestrial Ecosystem Research Network, Australia

More than 800 monitoring sites track status, history, trajectory and dynamics of ecosystems

​

3 of 10

Motivation for a Semantic Approach

  • Similar ecological data are collected at different spatial and program scales, with different field surveys
  • Inconsistent terminology makes cross-site and cross-program integration difficult.
  • Controlled vocabularies and ontologies provide a common semantic layer to harmonize datasets.
  • Machine-readable; Extensible information model for new surveys.

4 of 10

Harmonising Ecological Data with Vocabularies

Vocabularies as the primary harmonization tool

  • Platforms – instances modelled with SOSA (Sensor, Observation, Sample, and Actuator) and aligned with GCMD; platform types managed as SKOS concepts
  • Instruments – instances modelled with SSN; instrument types managed as SKOS concepts
  • Spatial regions – Australian bioregions (IBRA), ecoregions, states and territories
  • Spatial / temporal resolution, content type – represented in RDF and aligned with GCMD terms
  • Projects, organizations, people – modelled using schema.org patterns
  • Units of Measure (UoM) – drawn from the QUDT ontology
  • Observed properties / parameters – aligned with EnvThes and CF; supported by method/procedure vocabularies

5 of 10

Vocabulary Coverage and Source Integration

  • 20 concept schemes covering key ecological data aspects; largest scheme has 6526 parameter concepts (up to 6 levels).

​

  • 138 collections and 12,603 concepts describing platforms, instruments, properties, people, organisations.
  • 317 research platforms and 72 organisations represented as ontology instances.

​

  • External sources (GCMD, CF) brought in as local versions and linked with exactMatch / closeMatch to preserve interoperability and local control.

​

  • Editorial status flags (Draft, Published, Under Revision, Deprecated) support governance and provenance.

6 of 10

TERN Applications

  • Published vocabularies actively drive data discovery, harmonisation, and metadata authoring across TERN’s production applications.
  • Same vocabularies at submission and discovery time to remove ambiguity and make data FAIR.

7 of 10

Technical Architecture and Workflow

A curated workflow built on VocBench, GraphDB, SHACL, and automated synchronisation services delivers governed, versioned vocabularies to downstream applications.

  • GraphDB as RDF triple store and release repository (named graph versioning).

​

  • VocBench 3.0 for collaborative editing, deprecation, and quality control.

​

  • SHACL-based automated validation embedded in the editorial workflow.

​

  • Airflow DAGs to sync and refresh external sources (e.g. GCMD, CF).

​

  • Public exposure through TERN linked data viewer and publication to RVA.

8 of 10

Future steps

  • Integrate Large Language Models (LLMs) into our vocabulary management workflow to enhance semantic inference and search capabilities.

​

  • Publish all vocabularies to RVA to share them with the community.

​

  • Implement version control for all vocabularies, especially those that are updated periodically.

​

  • Release vocabulary change reports periodically to keep stakeholders informed.

9 of 10

Thank you!

Resources

​

​

​

​

10 of 10

We at TERN acknowledge the traditional owners and their custodianship of the lands on �which TERN operates. We pay our respects to their ancestors and their descendants, �who continue cultural and spiritual connections to country.

�

TERN is enabled by NCRIS.

Our work is a result of collaborative partnerships with many universities and institutions. 

To find out more please go to tern.org.au.

​