1 of 11

Recording uncertainty

2 of 11

Recording uncertainty

    • It summarizes the discussion in the emails

3 of 11

Types of uncertain data

  • Imprecision
    • E.g. the height of this book is 10cm ±2cm
  • Ambiguity
    • E.g. when the level of description does not match the level of query
  • Incompleteness
    • E.g. when some important information is not recorded
  • Uncertainty
    • Unsure about the truthiness of one or more statements (“epistemic uncertainty”)

  • Statistical error rate in a large set of statements �
    • (e.g. automatically extracted using software)

out of scope for this issue

out of scope for this issue

out of scope for this issue

out of scope for this issue

4 of 11

Proposed solutions (by SIG members)

  • Use of a .2 property to describe confidence/uncertainty
    • Suggested by Martin
  • Use of E13 Attribute Assignment
    • Suggested by Robert
  • Use a CRM extension (R1 Reliability Assessment, etc.)
    • Introduced by Franco Niccolucci and Sorin Hermon
  • Use CRMinf (properties 'J4 that' and 'J5 holds to be' and class I4 proposition set)

5 of 11

Main categories of solutions

  • At conceptual/ontological level
    • Use of specially designed classes and properties
      • E13 Attribute Assignment
      • R1 Reliability Assessment
      • I7 Belief Adoption

  • At implementation level
    • Use reification and relevant constructs
      • .2 properties
      • named graphs
      • singleton properties
      • RDF-star
    • Usually combined with a short vocabulary/model

  • Combined?
    • At implementation level, combined with provenance data by another actor (the ontological construct)
  • (?) Better when we want to describe someone's documented belief (e.g. statement X is probably true/false based on source Y)
    • We can query by uncertain value, or include uncertainty information in the query results
  • (?) Better when want to represent the maintainer’s uncertainty
    • We can query by uncertain value and filter out results that are below a certain threshold

6 of 11

How do we proceed?

7 of 11

A couple of important points

  • The uncertainty expressed by the maintainer of the knowledge base may differ from that of the creator of the statement ("partial belief adoption")
    • In this case, the different beliefs must be differentiated
    • Querying should only take into account the maintainers attribution (?)
    • A statement for which the maintainer has no opinion may only appear in an information object, and not be expanded in the KB as properties

  • The implementation as an attribute on an attribute existing in the knowledge base, be it with whatever construct (named graph, reification, ...), implies a possible reality assumed and expressed by the maintainer of the knowledge base
    • Querying such a KB will need filtering by uncertainty values, if the maintainers accept varying uncertainty values
    • This can be combined with a provenance by another actor (the ontological construct)

8 of 11

Categorization of different cases� that might need a different solution!

  1. How to express another's opinion (without taking a stance)
    • E.g. how to express Suetonius’ belief about Nero, how to express Tacitus’ belief about Nero

  • How to express someone’s (documented) opinion about another's opinion
    • E.g. how to express Francesca Bologna’s opinion about Nero, which supports that Suetonius belief is false while Tacitus’ belief is true?

  • How to express the maintainer’s own uncertainty/opinion about another's opinion
    • E.g. how to express my own uncertainty about Suetonius’ belief

  • How to express the maintainer’s uncertainty about her/his own statements
    • E.g. how to express my uncertainty about a specific piece of information

  • How to express the maintainer’s ‘unknown’ values
    • E.g. An observation was made but the observer could not recognize any known types (do we leave out such a statement, do we include it with some uncertainty value?)

9 of 11

Questions / Next steps

  • Qs:
    • Is there any other case?
    • How often do we encounter each case?

  • First next steps:
    • We need to gather real examples for each of the 5 cases (this might reveal details that we do not have yet in mind)
    • We also need query requirements (This will show if it makes sense to provide a very detailed (structured) representation of uncertainty and related provenance information)
    • Is there a good set of uncertainty values? (Computationally efficient are few, discrete values, in the way of fuzzy logic applications. Most probabilistic models have no objective base, except for observed error rates.)

  • Long-term objective:
    • Provide implementation recommendation for each case

10 of 11

Towards a good set of uncertainty values

  1. True (to the best of my knowledge)
  2. Probably True
  3. Unknown
  4. Probably False
  5. False (to the best of my knowledge)

Aim:

- To limit ambiguity

- To avoid fuzzy boundaries as much as possible

Any statement that we believe to be true may have a small chance of being wrong

Any statement that we believe to be false may have a small chance of being true

11 of 11

Thank you!