1 of 42

Retrieval Augmented Generation

Manish Gupta

gmanish@microsoft.com

Dec 2024

1

2 of 42

Retrieval Augmented Generation

2

3 of 42

RAG Details

  • Offline: Encode, Index (Non-parametric external memory)
  • Online: Encode, Retrieve, Combine, Generate
  • Solves Knowledge Cutoff Problem (Public Data)
    • No need to retrain LLM on new data
  • Access to Private Information
  • Reduce the chance of Hallucination 
    • By grounding an LLM on a set of external, verifiable facts.

3

4 of 42

Agenda

  • REALM (classification generation)
  • RAG (generation)
  • RETRO (scaling to trillion-sized index)
  • ATLAS (few-shot learning)
  • Internet-Augmented Generation (few-shot in-context learning)
  • TrieNLG (personalized QAC)
  • Multi-modal retrieval augmented generation
  • Summary

4

5 of 42

Agenda

  • REALM (classification generation)
  • RAG (generation)
  • RETRO (scaling to trillion-sized index)
  • ATLAS (few-shot learning)
  • Internet-Augmented Generation (few-shot in-context learning)
  • TrieNLG (personalized QAC)
  • Multi-modal retrieval augmented generation
  • Summary

5

6 of 42

Retrieval-Augmented Language Model (REALM)

  • REALM: Encoder+Retriever
    • Encoder: BERT based LM
    • Retriever: Another BERT + nearest neighbor lookup.
  • Knowledge corpus
    • 13M chunks from Enwiki docs
  • Find top k (k=5) matching docs efficiently and augment them to input for LM.

Guu, Kelvin, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. "Retrieval augmented language model pre-training." In ICML, pp. 3929-3938. PMLR, 2020.

6

7 of 42

Retrieval-Augmented Language Model (REALM)

  • Maximum Inner Product Search (MIPS)
    • Sim(question, docs).
    • Pre-compute doc embeddings.
    • Refresh periodically (every 500 steps)
  • Pretraining
    • MLM task: predict mask given sentence and retrieved docs. Unsupervised.
    • Mask salient spans (entities, dates)
    • Allow for null document.

Guu, Kelvin, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. "Retrieval augmented language model pre-training." In ICML, pp. 3929-3938. PMLR, 2020.

7

8 of 42

Retrieval-Augmented Language Model (REALM)

  • Finetuning
    • Open QA task: predict (start, end) spans given question and retrieved docs z.
  • Open QA datasets: NaturalQuestions, WebQuestions, CuratedTrec.
  • REALM > T5-11B while being 30x smaller.

Guu, Kelvin, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. "Retrieval augmented language model pre-training." In ICML, pp. 3929-3938. PMLR, 2020.

8

9 of 42

REALM Performance

Guu, Kelvin, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. "Retrieval augmented language model pre-training." In ICML, pp. 3929-3938. PMLR, 2020.

9

10 of 42

Agenda

  • REALM (classification generation)
  • RAG (generation)
  • RETRO (scaling to trillion-sized index)
  • ATLAS (few-shot learning)
  • Internet-Augmented Generation (few-shot in-context learning)
  • TrieNLG (personalized QAC)
  • Multi-modal retrieval augmented generation
  • Summary

10

11 of 42

Retrieval-Augmented Generation (RAG)

Lewis, Patrick, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler et al. "Retrieval-augmented generation for knowledge-intensive nlp tasks." NIPS (2020): 9459-9474.

  • Follow-up of REALM for NLG
  • Pre-trained Seq2Seq BART-Large combined with retriever
  • Retriever: Query Encoder + Document Index
    • MIPS to find top-k matching documents: FAISS with HNSW
    • BERT-base encoders
  • Jointly train retriever and generator.

11

12 of 42

Retrieval-Augmented Generation (RAG)

Lewis, Patrick, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler et al. "Retrieval-augmented generation for knowledge-intensive nlp tasks." NIPS (2020): 9459-9474.

  • RAG-Sequence: uses same document to predict each target token.
    • Generator produces the output sequence probability for each document, which are then marginalized.
  • RAG-Token: predict each target token based on a different document.
    • Generator can choose content from several documents when producing an answer.

12

13 of 42

How does RAG perform compared to BART and T5?

Lewis, Patrick, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler et al. "Retrieval-augmented generation for knowledge-intensive nlp tasks." NIPS (2020): 9459-9474.

https://ai.facebook.com/blog/retrieval-augmented-generation-streamlining-the-creation-of-intelligent-natural-language-processing-models/

  • Knowledge source: Wikipedia. 21M 100-word chunks.
  • Open-domain QA: Natural Questions, TriviaQA, WebQuestions, CuratedTrec
  • Abstractive QA (MSMARCO)
  • Jeopardy QG (SearchQA)

13

14 of 42

Agenda

  • REALM (classification generation)
  • RAG (generation)
  • RETRO (scaling to trillion-sized index)
  • ATLAS (few-shot learning)
  • Internet-Augmented Generation (few-shot in-context learning)
  • TrieNLG (personalized QAC)
  • Multi-modal retrieval augmented generation
  • Summary

14

15 of 42

What is Retrieval-Enhanced Transformer (RETRO)?

  • Combines a frozen BERT retriever, a differentiable encoder and a chunked cross-attention decoder.

Borgeaud, Sebastian, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche et al. "Improving language models by retrieving from trillions of tokens." In ICML, pp. 2206-2240. PMLR, 2022.

15

16 of 42

What is Retrieval-Enhanced Transformer (RETRO)?

  •  

Borgeaud, Sebastian, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche et al. "Improving language models by retrieving from trillions of tokens." In ICML, pp. 2206-2240. PMLR, 2022.

16

H

17 of 42

What is Retrieval-Enhanced Transformer (RETRO)?

  •  

Borgeaud, Sebastian, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche et al. "Improving language models by retrieving from trillions of tokens." In ICML, pp. 2206-2240. PMLR, 2022.

17

H

18 of 42

How does RETRO perform?

  •  

Borgeaud, Sebastian, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche et al. "Improving language models by retrieving from trillions of tokens." In ICML, pp. 2206-2240. PMLR, 2022.

18

19 of 42

Agenda

  • REALM (classification generation)
  • RAG (generation)
  • RETRO (scaling to trillion-sized index)
  • ATLAS (few-shot learning)
  • Internet-Augmented Generation (few-shot in-context learning)
  • TrieNLG (personalized QAC)
  • Multi-modal retrieval augmented generation
  • Summary

19

20 of 42

ATLAS architecture and training

20

21 of 42

ATLAS architecture and training

  •  

21

22 of 42

ATLAS architecture and training

  • Contriever (BiEncoder) Retriever
    • Index updates
      • Full update
      • Re-rank
      • Query-side finetuning with fixed doc index.
    • Query-side fine-tuning for experiments with small numbers of examples, and standard fine-tuning for larger datasets.
  • LM: T5, fusion-in-decoder.

22

23 of 42

How does Atlas perform?

  • Knowledge-Intensive Language Tasks (KILT): fact checking, question answering, dialog generation, entity linking and slot-filling
  • Atlas-11B achieves 42.4% acc on NQ using 64 training examples, outperforming PaLM 540B by ~3 points
  • Atlas-11B achieves SOTA 60.4% in a full-dataset setting.

Izacard, Gautier, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. "Few-shot learning with retrieval augmented language models." arXiv:2208.03299 (2022).

23

Question Answering

24 of 42

How does Atlas perform?

  • Benchmarks
    • Massively-Multitask Language Understanding (MMLU): 57 MCQ datasets.
    • FEVER

Izacard, Gautier, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. "Few-shot learning with retrieval augmented language models." arXiv:2208.03299 (2022).

24

MMLU

FEVER

25 of 42

Agenda

  • REALM (classification generation)
  • RAG (generation)
  • RETRO (scaling to trillion-sized index)
  • ATLAS (few-shot learning)
  • Internet-Augmented Generation (few-shot in-context learning)
  • TrieNLG (personalized QAC)
  • Multi-modal retrieval augmented generation
  • Summary

25

26 of 42

Few-shot prompting for Internet-augmented LMs

Lazaridou, Angeliki, Elena Gribovskaya, Wojciech Stokowiec, and Nikolai Grigorev. "Internet-augmented language models through few-shot prompting for open-domain question answering." arXiv:2203.05115 (2022).

  • Retrieve
  • Prompt
  • Rerank

26

27 of 42

Few-shot prompting for Internet-augmented LMs

  • Retrieve: Given a question q, retrieve a set of 20 relevant documents D from the web using Google
    • 6 sentence chunks
    • Embed q and the paragraphs using TF-IDF and using cosine similarity to produce a (ranked) list of evidence paragraphs P. Take n=50.
  • Prompt: use the retrieved evidence to condition the LM through few-shot prompting
    • Evidence: ... Question: ... Answer: ...
    • k = 15 shot prompts, where we use the gold evidence documents as evidence
    • Generate 200 answers.

Lazaridou, Angeliki, Elena Gribovskaya, Wojciech Stokowiec, and Nikolai Grigorev. "Internet-augmented language models through few-shot prompting for open-domain question answering." arXiv:2203.05115 (2022).

27

28 of 42

Few-shot prompting for Internet-augmented LMs

  •  

Lazaridou, Angeliki, Elena Gribovskaya, Wojciech Stokowiec, and Nikolai Grigorev. "Internet-augmented language models through few-shot prompting for open-domain question answering." arXiv:2203.05115 (2022).

28

29 of 42

Results

  • Datasets:
    • QA: single-hop NQ, multi-hop HOTPOTQA and STRATEGYQA
    • Fact-checking multi-hop dataset FEVER
  • Language generation (NQ, HOTPOTQA) and classification (2-way for STRATEGYQA, 3-way for FEVER)
  • LM: GOPHER 44M, 117M, 400M, 1B, 7B, 280B trained on MassiveText 300B tokens.
  • CB=Closed Book baseline.

Lazaridou, Angeliki, Elena Gribovskaya, Wojciech Stokowiec, and Nikolai Grigorev. "Internet-augmented language models through few-shot prompting for open-domain question answering." arXiv:2203.05115 (2022).

29

Results on 4 question answering datasets using the GOPHER-280B model.

30 of 42

Agenda

  • REALM (classification generation)
  • RAG (generation)
  • RETRO (scaling to trillion-sized index)
  • ATLAS (few-shot learning)
  • Internet-Augmented Generation (few-shot in-context learning)
  • TrieNLG (personalized QAC)
  • Multi-modal retrieval augmented generation
  • Summary

30

31 of 42

QAC and TrieNLG

  • MPC: Most Popular Completion
    • Capture popularity
  • NLG: Capture personalization using prev session queries.
  • Challenge: Short and unseen prefixes.
  • TrieNLG
    • Retrieval
      • Extracts up to top-m most popular completions from the 1B suggestions trie
      • For unseen prefixes, use suffix word ngrams trie
    • Generator: BART
      • Input: session queries, top m MPC suggestions, current prefix.
      • Trained to maximize the probability of ground truth token sequence with MLE.

Kaushal Maurya, Maunendra Sankar Desarkar, Manish Gupta, Puneet Agrawal. TrieNLG: Trie Context Augmentation to Improve Personalized Query Auto-Completion for Short and Unseen Prefixes. ECML-PKDD 2023.

31

32 of 42

QAC and TrieNLG

Kaushal Maurya, Maunendra Sankar Desarkar, Manish Gupta, Puneet Agrawal. TrieNLG: Trie Context Augmentation to Improve Personalized Query Auto-Completion for Short and Unseen Prefixes. ECML-PKDD 2023.

  • Best results when 3 trie suggestions are augmented.
  • Trie lookups are very cheap compared to BART-based suggestion generation.

32

Bing dataset (improvements wrt MPCTrain+MPCSynth)

33 of 42

Agenda

  • REALM (classification generation)
  • RAG (generation)
  • RETRO (scaling to trillion-sized index)
  • ATLAS (few-shot learning)
  • Internet-Augmented Generation (few-shot in-context learning)
  • TrieNLG (personalized QAC)
  • Multi-modal retrieval augmented generation
  • Summary

33

34 of 42

KAT: Knowledge Augmented Transformer for Vision-and-Language

Gui, Liangke, Borui Wang, Qiuyuan Huang, Alex Hauptmann, Yonatan Bisk, and Jianfeng Gao. "Kat: A knowledge augmented transformer for vision-and-language." arXiv:2112.08614 (2021).

34

35 of 42

How does KAT perform?

Gui, Liangke, Borui Wang, Qiuyuan Huang, Alex Hauptmann, Yonatan Bisk, and Jianfeng Gao. "Kat: A knowledge augmented transformer for vision-and-language." arXiv:2112.08614 (2021).

35

36 of 42

REVEAL: Retrieval-Augmentation with Multi-Source Multimodal Knowledge Memory

Hu, Ziniu, Ahmet Iscen, Chen Sun, Zirui Wang, Kai-Wei Chang, Yizhou Sun, Cordelia Schmid, David A. Ross, and Alireza Fathi. "Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge memory." In CVPR, pp. 23369-23379. 2023.

36

37 of 42

Results

Hu, Ziniu, Ahmet Iscen, Chen Sun, Zirui Wang, Kai-Wei Chang, Yizhou Sun, Cordelia Schmid, David A. Ross, and Alireza Fathi. "Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge memory." In CVPR, pp. 23369-23379. 2023.

37

REVEAL can use knowledge from different sources to correctly answer the question.

38 of 42

Retrieval-Augmented Multimodal Language Modeling

38

39 of 42

Retrieval-Augmented Multimodal Language Modeling

39

Knowledge-intensive multimodal generation

40 of 42

Summary

  • REALM: LM + knowledge retriever. Supported by MIPS. Great results for OpenQA task.
  • RAG: Does NLG. Better than standard NLG models like BART and T5. Combines BART-Large with non-parametric BERT-base retriever.
  • RETRO: Retrieval-Enhanced Transformer. Combines a frozen BERT retriever, a differentiable encoder and a chunked cross-attention decoder. Trillion sized index. Interleaved RETRO-blocks with Chunked Cross attention.
  • ATLAS: 770M, 3B, 11B checkpoints. Strong few-shot learning capabilities. Atlas-11B > 540B PaLM, which required 50x more pre-training compute.
  • Internet-Augmented Generation: Retrieve, Prompt, Rerank. PoE outperforms consistently. Searching the Internet is worth >273 billion parameters.
  • TrieNLG: Prepend prefix with trie completions for QAC.
  • Multimodal RAG: KAT and REVEAL

40

41 of 42

Problems (specifically for search grounding)

  • How do you compress retrieved documents?
  • How do you scale RAG to 100s of billions of documents?
  • How do you do RAG on knowledge graphs and spatio-temporal data?
  • How do you reduce latency even when doing multi-vector dense retrieval?
  • How do you decide budget across different retrieval document types (like web documents, structured json, feeds data, images, videos)?
  • RAG on videos?

41

42 of 42

Thanks!

42