1 of 71

CSCI-SHU 376: Natural Language Processing

Hua Shen

2026-03-19

Spring 2026

Lecture 13: Retrieval-Augmented Language Model

2 of 71

Today’s Plan

  • Retrieval-augmented Language Model
  • Retrieval: Embedding Learning

3 of 71

Retrieval-based Language Models (RALM)

Inference

  • Retrieval-based LM = Retrieval + LMs (or commonly referred as RAG)
  • It retrieves from an external datastore

3

4 of 71

Benefit of RALM #1: Hallucinations

Inference

  • Hallucination: The generation of content that is not grounded in the input data or external knowledge, often including false information that appears plausible.

“The 0.3 cm x 0.4 cm x 0.3 cm oval mass in the left breast at 10 o'clock posterior depth likely represents a complicated cyst.

Language Model Hallucinates, should be 7 o’clock

4

5 of 71

Inference

  • Hallucination is not always a bad thing!
  • Encourage creativity!

It is bad under critical domains: medical, law, etc.

Benefit of RALM #1: Hallucinations

5

6 of 71

Inference

  • LM’s parametric knowledge gets outdated quickly

Benefit of RALM #2: Adaptations

6

7 of 71

Inference

  • We can easily swap datastores of RALMs for new data distributions

Benefit of RALM #2: Adaptations

7

8 of 71

Inference

  • Provide references (e.g., citations) as attributions
  • Increase user trust

Benefit of RALM #3: Attributions

8

9 of 71

Inference

  • Incorporate / remove high-risk data dynamically at inference

Benefit of RALM #4: Private Data

9

10 of 71

RALM has been widely used

10

11 of 71

Tool-augmented Language Model

  • RALM is a special Tool-augmented LM

11

12 of 71

Tool-augmented Language Model

  • RALM is a special Tool-augmented LM

12

13 of 71

History of RALM

  • RALM is initially studied in open-domain QA setting

13

14 of 71

History of RALM

  • RALM is initially studied in open-domain QA setting

14

15 of 71

History of RALM

  • RALM is initially studied in open-domain QA setting

15

16 of 71

History of RALM

  • RALM is initially studied in open-domain QA setting

16

17 of 71

History of RALM

  • Current: LLM-based RAG systems for diverse use cases

17

18 of 71

Today’s Plan

  • Retrieval-augmented Language Model
  • Retrieval: Embedding Learning

19 of 71

Inference

  • Retrieval in Open-domain QA: Find relevant information that (hopefully) contains the final answer

Information Retrieval

19

20 of 71

Inference

  • A TF-IDF weighted term vector model over unigrams / bigrams
  • This retriever is not trainable

Sparse Retriever

20

21 of 71

Inference

  • Dense representations have never been better than sparse representations before 2019

Sparse vs dense representations

21

22 of 71

Inference

  • Pre-training model (e.g., BERT) matters!
  • Large enough labelled data (e.g., 82M query-doc pairs from user clicks)
  • Better techniques and tools to support fast maximum inner product search (MIPS)

Why dense retrieval now?

22

23 of 71

Inference

  • Encode data into an embedding vector
    • Fine-tune LLMs

  • Key idea: Capture semantic relationships through distances in embedding space

Dense Retrieval: Embedding learning

23

24 of 71

Inference

Dense Retrieval: Embedding learning

24

25 of 71

Inference

Dense Passage Retrieval (DPR)

  • Key Contribution: you can train a dense retriever only from a small number of Q/A pairs

25

26 of 71

Inference

Dense Passage Retrieval (DPR)

  • Selecting Positives:
    • Provided in the datasets
    • Passages of high BM25 scores that contain the answer string

  • Negatives ??
    • Corpus size is huge
    • 99.9999% are trivially irrelevant
    • All about finding hard negatives!

26

27 of 71

Inference

Negative Samples Selection - DPR

  • Random passages from the corpus
  • Passages of high BM25 scores that do not contain the answer string
  • In batch negatives: positive passages of other questions in the mini-batch

Issues:

  • Often negatives from sparse retrieval are trivial for dense retrieval
  • Weaker generalizability empirically

27

28 of 71

Inference

Negative Samples Selection - ANCE

  • Approximate Nearest Neighbor Negative Contrastive Learning
  • Sampling negatives from trained model itself
  • Periodically refresh the dense retrieval index to keep updated

28

29 of 71

Inference

Error Analysis

  • BM25 and ANCE only agree on 20% of their Top 100 rankings
  • But both find relevant document in Top 3

29

30 of 71

Inference

Mismatch between LLM and Embeddings

  • Intrinsic Difference between LLM’s capacity and Embedding model
  • Instruction Following in Embeddings
  • Current Research: Generative Representational Instruction Tuning

30

31 of 71

Performance on MTEB

  • Ongoing Research: How to distil LLM capacities into Embeddings
  • MTEB: Massive Text Embedding Benchmark

31

32 of 71

Today’s Plan

  • Retrieval-augmented LM Architectures
  • Knowledge Conflict
  • Ongoing Research: Search-R1

33 of 71

Diverse architectures of RALM

Inference

  • Input augmentation
  • Intermediate incorporation
  • Output incorporation

33

34 of 71

Diverse architectures of RALM

Inference

  • Input augmentation
  • Intermediate incorporation
  • Output incorporation

34

35 of 71

RALM: Input Augmentation

Inference

  • Augment the inputs with retrieved context
  • E.g., RAG, REALM, In-context RALM etc

35

36 of 71

REALM: Augmenting input space of LMs

Inference

36

37 of 71

REALM: Augmenting input space of LMs

Inference

13M Wikipedia Passages

37

38 of 71

REALM: Augmenting input space of LMs

Inference

38

39 of 71

REALM: Augmenting input space of LMs

Inference

39

40 of 71

Retrieval augmented generation (RAG)

  • Original paper of RAG
  • RAG combines a trained retriever and autoregressive BART as generator

40

41 of 71

Results

Inference

  • RAG outperforms REALM and other baselines on Open-domain QA

41

42 of 71

In-context Retrieval-augmented LMs

Inference

  • Often referred as “RAG” nowadays

42

43 of 71

Pros and Cons of Input Augmentation

Inference

  • Pros:
    • (Probably) No need to fine-tune / without training

  • Cons:
    • LLMs (still) struggle on long-context

43

44 of 71

Intermediate incorporation

Inference

  • Incorporate retrieved context in intermediate spaces of Transformers
  • E.g., RETRO

44

45 of 71

RETRO: Incorporate context in intermediate layers

Inference

  • Incorporate retrieved context in intermediate spaces of Transformers
  • E.g., RETRO

45

46 of 71

RETRO: Incorporate context in intermediate layers

Inference

  • Given the input sequence, first retrieves a set of relevant documents

46

47 of 71

RETRO: Incorporate context in intermediate layers

Inference

  • Use cross-attention to generate retrieved context-aware representations

47

48 of 71

RETRO: Incorporate context in intermediate layers

Inference

  • Use cross-attention to generate retrieved context-aware representations

48

49 of 71

RETRO: Incorporate context in intermediate layers

Inference

  • Concatenate all of the CA output

49

50 of 71

RETRO: Incorporate context in intermediate layers

Inference

  • RETRO shows impressive performance improvements on upstream tasks

50

51 of 71

Pros and Cons of Intermediate Augmentation

Inference

  • Pros:
    • Good empirically results

  • Cons:
    • Require modification of underlying LMs
    • Therefore, hard to take advantage of recent LMs

51

52 of 71

Output interpolation

Inference

  • Interpolate output token probabilities with retrieved non-parametric distributions
  • E.g., KNN-LM

52

53 of 71

KNN-LM: directly interpolate token distributions

Inference

  • A LM predicts parametric distributions of next token

53

54 of 71

KNN-LM: directly interpolate token distributions

Inference

  • A LM predicts parametric distributions of next token
  • KNN-LM computes non-parametric distributions

54

55 of 71

KNN-LM: directly interpolate token distributions

Inference

  • Interpolate two token distributions, adjusting the balance using a hyperparameter

55

56 of 71

KNN-LM: directly interpolate token distributions

Inference

  • KNN-LM outperforms much larger parametric LMs by large margin

56

57 of 71

Pros and Cons of Output Augmentation

Inference

  • Pros:
    • Explicit control between parametric and non-parametric distribution

  • Cons:
    • Difficult to scale to large retrieval corpora
    • Empirically show limited progress outside of language modeling tasks

57

58 of 71

Today’s Plan

  • Retrieval-augmented LM Architectures
  • Knowledge Conflict
  • Ongoing Research: Search-R1

59 of 71

Inference

  • LM’s parametric knowledge gets outdated quickly

Recap: Benefit of RALM, Adaptations

59

60 of 71

Inference

  • Goal: Update LLM with up-to-date knowledge

Update LLM knowledge

60

61 of 71

Inference

  • Direct parameter editing
  • Add extra trainable editable parameters

One Attempt: Knowledge Editing

61

62 of 71

Inference

  • Evaluation: edited facts and unchanged facts

  • Ripple Effects

  • High-level take-away: Not working yet!

Issues with Knowledge Editing

62

63 of 71

Inference

  • Retrieve relevant context first!

  • Knowledge conflict
    • Parametric Knowledge
    • Contextual Knowledge

Alternative Approach: RALM

63

64 of 71

Inference

  • Identify knowledge conflicts
  • Pinpoint conflicting information segments
  • Provide distinct answers

Knowledge Conflict

64

65 of 71

Inference

  • Binary Classification
  • Prompting is reasonably good
  • Representations?

Knowledge Conflict: Identification

65

66 of 71

Inference

  • Find relevant spans
  • LLM somewhat struggles

Knowledge Conflict: Localization

66

67 of 71

Inference

  • Trust the parametric knowledge?
  • Trust the context?
  • Controllable generation!

Knowledge Conflict: Generation

67

68 of 71

Today’s Plan

  • Retrieval-augmented LM Architectures
  • Knowledge Conflict
  • Ongoing Research: Search-R1

69 of 71

Answer

Which city state was OpenAI founded?

OpenAI was founded in

San Francisco in late 2015…

Evidence

Multi-Step

Retriever

Multi-hop Reader

California

San Francisco is the … center of Northern California in the United States

Connected

Multi-hop Question Answering

California

69

70 of 71

Inference

RALM with Large reasoning models

  • Reasoning Model (e.g., R1) as planner
  • Advanced reasoning behaviors?
    • Reflection, revision etc?

70

71 of 71

Inference

Task: Reasoning-intensive Retrieval

71