1 of 20

Reasoning-driven Question Answering

Daniel Khashabi

​

​

​

NYU

Sept, 2018

@DanielKhashabi

2 of 20

Exams

2

Q: Which physical structure would best help a bear to�survive a winter in New York State?

A: (A) big ears (B) black nose (C) thick fur (D) brown eyes

​

Standardized science exams (Clark et al, 2015):

  • Simple language; kids can solve them well, but they need to have the ability use the knowledge and abstract over it.

Biology exams (Berant et al, 2014):

  • Technical terms and answer not easy to find.
  • Requires understanding of complex relations.

Q: What does meiosis directly produce?

(A) Gametes (B) Haploid cells

P: … Meiosis produces not gametes but haploid cells that then divide by mitosis and give rise to either unicellular descendants or a haploid multicellular adult organism. Subsequently, the haploid organism carries out further mitoses, producing the cells that develop into gametes.

P: … Polar bears, saved from the bitter cold by their thick fur coats, are among the animals in danger …

3 of 20

Linguistic variability

3

Which physical structure would best help a bear to survive a winter?

(A) big ears (B) black nose (C) thick fur (D) brown eyes

A given “meaning” can be phrased in many surface forms!

Thick fur helps a bear survive a winter.

A thick coat of white fur helps bears survive in these cold latitudes.

Polar bears, saved from the bitter cold by their thick fur coats, are among the animals in danger of extinction because of the global warming and human activities.

4 of 20

QA is a language understanding problem!

4

Which physical structure would best help a bear to survive a winter?

(A) big ears (B) black nose (C) thick fur (D) brown eyes

Polar bears, saved from the bitter cold by their thick fur coats, are among the animals in danger of extinction because of the global warming and human activities.

QA is fundamentally an NLU problem

preposition

comma

A single abstraction is not enough

verb

5 of 20

High-level view

5

Question Answering

Question Answering

Semantic Abstractions

Global Reasoning

as Global Reasoning

over Semantic Abstractions

6 of 20

Collections of semantic graphs

Create a unified representation of families of graphs

    • predicate-argument, trees, clusters, sequences

6

A single representation is not enough to capture the complexity of language

e.g named-entities

e.g co-reference

e.g semantic role labeling �(verb, preposition, comma)

e.g dependency parse

e.g tables

Our representation has nothing to do with the QA task. It reflects our understanding of the language

TableILP: IJCAI’16

- Surface word

- Label, e.g. subj.

- W2V representation

…

5

Consequently, we expect these representations to be useful for a range of tasks

7 of 20

Reasoning With a Meaning Representation

  • Augmented Graph is the graph which contains potential alignments between elements of any two graphs

7

QA Reasoning formulated as finding “best” explanation – subgraph connecting Q to A via P

Question Instance

​

Question

Paragraph

Answer

Edges reflect similarity / entailment

This is a realization of

abductive reasoning!

​

​

​

​

​

​

​

​

​

(Incomplete)

Observations

Best explanation (maybe true)

8 of 20

Example subgraph

8

Question Instance

​

Question

Paragraph

Answer

(Irrelevant edges and graphs are dropped for simplicity)

9 of 20

SemanticILP, some details.

Translate QA into a search for an optimal subgraph

​

9

Formulate as Integer Linear Program (ILP) optimization

    • Solution points to the best supported answer

Objective: Capture what’s a valid reasoning, what’s preferred

    • Preferences e.g.
      • Use sentences nearby
      • If using a pred-arg graph, give priority to the subject

Constraint: Incorporate global and local constraints

    • Global e.g.
      • Have ends in question and paragraph
      • Connected graph
    • Local e.g.
      • If using a pred-arg graphs,
        • use at least predicate and argument, or
        • use at least two arguments

10 of 20

Evaluation: notable baselines

  • IR (Clark et al, AAAI’15)
    • Information retrieval baseline (Lucene)
    • Using 280 GB of plain text

​

Thick white fur is an animal adaptation most needed for the climate in which biome?

(A) deserts (B) taiga (C) deciduous forest (D) tundra

  • TupleINF (Khot et al, ACL’17)
    • Inference over independent rows
    • Auto-generated short triples
    • And type-constrained rules
  • BiDaF (Seo et al, ICLR‘16)
    • Neural model: attention & LSTM
    • Extractive, i.e select a contiguous phrase in a given paragraph

10

Type constrained rules:�(X, helps in, Y), (Z, has, Y) => (X, helps in, Z)

Thick fur

helps in

cold winter

Tundra biome

has

cold winter

helps in

We compare with the best baseline on each domain.

However we use one version of our systems across all the datasets.

11 of 20

Results #1: Science Questions

11

(exam scores, shown as a percentage)

Higher is better

[AAAI, 2018]

12 of 20

Results #2: Biology Questions

12

A single system tested on different datasets.

More experiments

in the paper!

Using additinal supervision

[AAAI, 2018]

13 of 20

A Comment on Robustness

13

How robust are approaches to simple question perturbations�that would typically make the question easier for a human?

​

  • E.g., Replace incorrect answers with arbitrary co-occurring terms

​

[IJCAI’16]

In New York State, the longest period of daylight�occurs during which month?

(A) March (B) June (C) September (D) December

[Jia&Liang,EMNLP’17, ...]

years

history

eastern

14 of 20

There must be some limits

There must be limits to multi-step reasoning

  • Inaccuracy of grounding and sparsity of information
  • More noise accumulate with each step

14

Reasoning problems that require longer reasoning tend to be harder

Assembling chains with more than 3-4 facts suffers from "semantic drift" (Fried et al, 2015, Jansen et al 2017)

Figure credit: �Peter Jansen

15 of 20

The impossibility of long-range reasoning

15

H1

(Conceptualization)

Language as communication channel�(with symbolic information)

Observations

H2

at least one of the hypotheses require d-step connectivity within a graph with n nodes

Theorem (informal)

If d∈Ω(log n)

​

no algorithm can confidently distinguish between H1 and H2.

and “sufficient” noise

[In Submission]

a small number

16 of 20

Lessons from “theoretical limitations”

  • This might not be a big problem (for now):
  • Average humans (hence many of the datasets) are limited.
  • To go beyond
  • Strong correlation between working memory capacity and reasoning ability (Kyllonen and Christal, 1990)
  • Humans, on average, can maintain few (4~7) "facts" simultaneously in their active memory (Cowan2001)
  • Being able to abstract and ground seems to be a key element.

16

There might be fundamental barriers for a class of reasoning systems.

17 of 20

Pushing it further

17

  • We need reading comprehension playground which requires deeper “reasoning”

​

  • We propose a reading comprehension challenge that require more than one sentence in order to answer

​

​

  • Does not restrict us to a narrow class of “reasoning” phenomena

[NAACL, 2018]

https://cogcomp.org/multirc

18 of 20

Wrap up

  • Reasoning over language requires dealing with a diverse set of semantic phenomena.
  • There are limits to “naive” chaining.
  • To go beyond, one probably need to use better grounding and abstractions.

18

  • Semantic variability ⇒ collection of semantic abstractions that are linguistically informed
  • We decoupled “reasoning for QA” from “abstraction”
  • Strong performances on two domains simultaneously

System design

Limits of reasoning

19 of 20

Acknowledgment

19

Dan Roth (UPenn)

Tushar Khot (AI2)

Ashish Sabharwal (AI2)

Snigdha Chaturvedi (UCSC)

Michael Roth �(Saarland Univ)

Shyam Upadhyay �(Uepnn)

Peter Clark �(AI2)

Oren Etzioni �(AI2)

Erfan Sadeqi Azer (Indiana U)

20 of 20

Questions?

That’s it folks

20