1 of 53

CS 162: Natural Language Processing

Saadia Gabriel

Lecture 10:

Information Retrieval & Extraction Part 1

2 of 53

Announcements

  • Bruin Learn is back!
  • Homework 3 will out later today (due 5/29).
  • Student presentations start today, make sure to sign up and check guidelines.

3 of 53

Q & A

Before large-scale language modeling, BERT was better for discriminative tasks like MCQA and text classification. Now we can have massive generative models with large-scale pretraining that can perform most tasks well and are more flexible than a bidirectional model. BERT still often used when a lightweight encoder is needed.

The numbers that define the weights of pretrained models update based on information from the training data, so they implicitly encode semantics, world knowledge, etc.

4 of 53

Q & A

It depends on the complexity of the language task and test data, some are easier to model based on presence of keywords (e.g. sentiment analysis).

Yes! Next lecture we’ll learn about hidden Markov models (HMMs) and conditional random fields (CRFs), which are particularly useful for sequence tagging problems.

5 of 53

Goals of This Lecture

  • NLP Motivation for Information Retrieval
  • Information Extraction
  • Background on Information Retrieval

6 of 53

Recap:

LLMs make factual errors known as “hallucinations”

and they do this often.

7 of 53

Recap: Prompting Solution

8 of 53

Later on we’ll talk about retrieval-augmented generation as another method to improve accuracy!

9 of 53

Today is later…

10 of 53

Origins of Information Retrieval

- Manning, Schutze and Raghavan (2008)

- Manning, Schutze and Raghavan (2008)

Term-document boolean incidence matrix

Query: Which Shakespeare plays contain the words Brutus AND Caesar AND NOT Calpurnia?

Logical bitwise operations over word vectors:

Complement for NOT

11 of 53

Types of Retrieval Methods

Dense vector retrieval

Sparse retrieval, e.g. BM25

Hybrid retrieval

Look for semantically similar embeddings

Bag-of-words approach that rank document relevance based on keyword presence

A combination of dense and sparse approaches

12 of 53

What type of retrieval is the boolean example?

13 of 53

Types of Retrieval Methods

Dense vector retrieval

Sparse retrieval, e.g. BM25

Hybrid retrieval

Look for semantically similar embeddings

Bag-of-words approach that rank document relevance based on keyword presence

A combination of dense and sparse approaches

14 of 53

Retrieval-augmented Generation

Figure courtesy of Holt Skinner, Google I/O

15 of 53

Retrieval-augmented Generation

Query prompt

LLM

Knowledge cut-off: January 2025

Who is the president of UC?

🧑

Output

Michael Drake

As of August 2025:

16 of 53

Retrieval-augmented Generation

Query prompt

LLM

Knowledge cut-off: January 2025

Who is the president of UC?

🧑

Output

Michael Drake

As of August 2025:

We can fix this without making any changes to the underlying model!

17 of 53

Query prompt

Retrieval-augmented Generation

We can fix this without making any changes to the underlying model!

LLM

Who is the president of UC?

🧑

Knowledge cut-off: January 2025

Document Corpus

Embed query

Embed docs

Dense vector retrieval

Retrieve relevant doc based on a metric like cosine similarity

Output

James Milliken,

according to source

Example of retrieved passages from your text

18 of 53

Retrieval-Augmented Generation

Slides courtesy of Graham Neubig

19 of 53

Retrieval-Augmented Generation

Recall there are many types of retrieval methods

Slides courtesy of Graham Neubig

20 of 53

Retrieval-Augmented Generation

Dense retrieval

Slides courtesy of Graham Neubig

21 of 53

Retrieval-Augmented Generation

Learning retrieval-oriented embeddings

Slides courtesy of Graham Neubig

22 of 53

Retrieval-Augmented Generation

Simple retrieval + reading approach

Slides courtesy of Graham Neubig

23 of 53

Retrieval-Augmented Generation

Slides courtesy of Graham Neubig

24 of 53

Agentic RAG

RAG frameworks are now more adaptive due to integration with Agentic AI frameworks

LLM agents can work together to route queries to specialized external data sources

Agents can also refine queries and generate plans for how to optimally use retrieved evidence

25 of 53

QA Recap

Courtesy of Diyi Yang

Protosynthex system, which retrieved candidate answers based on term overlap

Dependency parsers were then used to see which answer mostly closely matched the question.

26 of 53

Textual QA

(Reading Comprehension)

Courtesy of Diyi Yang

27 of 53

Textual QA

(Reading Comprehension)

Courtesy of Diyi Yang

28 of 53

Open-domain QA

V

V

Courtesy of Diyi Yang

29 of 53

Courtesy of Diyi Yang

Where RAG Fits In

30 of 53

Where RAG Fits In

Courtesy of Diyi Yang

However, we still face challenges of how many documents to retrieve (context constraints) and inaccurate or hallucinated citations.

31 of 53

Information Extraction

32 of 53

A PubMed Article

What is this saying?

33 of 53

A PubMed Article

There is a lot of jargon.

Perhaps if we can automatically extract these domain-specific terms,

we can look them up and find an easier to understand description.

34 of 53

A PubMed Article

We refer to these technical terms as “entities.”

35 of 53

We are transitioning from unstructured knowledge…

36 of 53

We can assign types to the entities, like organizations (ORG)

We can also classify relationships between entities, like PART-WHOLE

… to structured knowledge

37 of 53

Structured Information Comes in Many Forms

Knowledge Graph

Database Table

38 of 53

We Can Use Entity Extraction for Classification

39 of 53

We Can Use Entity Extraction for Classification

40 of 53

Named Entities

41 of 53

Named Entity Recognition

42 of 53

Example of NER Output

One commonly used NER recognizer is Stanford NER (https://nlp.stanford.edu/software/CRF-NER.html), which is based on a CRF (another probabilistic model we’ll learn about these later).

43 of 53

Example Applications

44 of 53

Challenges of NER

What do you think the ambiguities are?

Is Washington a person, a sports team, a place, a political entity?

How do we know to segment “James Burroughs” as a single named entity?

45 of 53

Supervised Learning Setting

46 of 53

BIO Tagging

47 of 53

NER with BIO Tagging

https://www.google.com/url?sa=i&url=https%3A%2F%2Fwww.kaggle.com%2Fcode%2Fdevashishprasad%2Fbio-labeled-dataset%2Fnotebook&psig=AOvVaw2ImnUI01nLjBrguSeXuqba&ust=1747159874974000&source=images&cd=vfe&opi=89978449&ved=0CBcQjhxqFwoTCPDEqvHDno0DFQAAAAAdAAAAABAb

48 of 53

Hands-on Exercise

Pretrained language model for NER

49 of 53

https://cogcomp.seas.upenn.edu/page/demo_view/NEREnglish

Let’s check out a NER demo

What are the NER tags for the following sentences:

1 ) Ta-Nehisi Coates’ first book was The Beautiful Struggle.

2 ) On my way to Apple I dropped my iPhone in an apple pie.

50 of 53

51 of 53

52 of 53

Next Time

  1. Probabilistic models for sequence tagging (so what are CRFs???)
  1. Parts of speech (POS tagging)

53 of 53

Student Presentations