1 of 13

Frugal Prompting for Dialog Models

Bishal Santra1*, Sakya Basak2*, Abhinandan De1, Manish Gupta2, Pawan Goyal1

{bishal.santra@,abhinandan0316@,pawang@cse}.iitkgp.ac.in; {sakya.basak,gmanish}@microsoft.com

1IIT-Kharagpur, 2Microsoft

1

* Equal contribution

2 of 13

Huge inference costs of LLMs

  • In-context learning based LLMs have changed NLP.
  • GPT-3, CODEX, LaMDA, PaLM are closed source.
  • Billing by LLM-inference APIs depends on input and output size.
  • How can we design prompts to increase LLM perf but reduce prompt size, in an ICL setup?
  • Contributions
    • Effectiveness of various ICL models and input formats for dialog modeling.
    • UID (Usable Information Density) to capture tradeoff between accuracy and length for various (input format, ICL model) combinations.
    • Expts with 2 datasets (MSC and TC) and 4 ICL models (GPT3, FLAN-T5, T0, Tk-Instruct)

Bishal Santra, Sakya Basak, Abhinandan De, Manish Gupta, Pawan Goyal. Frugal Prompting for Dialog Models. EMNLP Findings, 2023.

2

3 of 13

Related Work

  • LLMs for dialog modeling
    • DialoGPT, Plato, Blenderbot-3 (175B), Meena and LaMDA (137B)
      • Compute Intensive to finetune.
    • OpenAI textdavinci-003 trained using RLHF
    • In-context learning (ICL)
      • T0, FLAN, Tk-Instruct
      • Increased inference costs due to large prompts sizes

Bishal Santra, Sakya Basak, Abhinandan De, Manish Gupta, Pawan Goyal. Frugal Prompting for Dialog Models. EMNLP Findings, 2023.

  • Optimize computation for LLMs
    • Environmental impact in terms of CO2 emissions.
    • Can we make inference step of a transformer efficient?
    • Model distillation-based methods
    • Efficient transformer architectures
      • Reformer, Linformer, BigBird, Longformer, etc.

3

4 of 13

Prompt Ingredients for Dialog Systems

  • Task Instruction
    • Explain the task of a dialog response generation model. Assign a system-role (chat assistant).
  • Dialog Context
    • Dialog history
    • Background Information (BI)
      • Persona: fictional representation of a user
      • Knowledge sections: short paragraphs from Wikipedia, Reddit, and Washington Post related to topic of conv.
  • Person1’s latest utterance
  • Exemplars: zero-shot vs few-shot

Bishal Santra, Sakya Basak, Abhinandan De, Manish Gupta, Pawan Goyal. Frugal Prompting for Dialog Models. EMNLP Findings, 2023.

4

5 of 13

Manual versus Perplexity Prompts

  • Manually Designed Prompts
    • Avoid repetitive, dull responses and maintain consistency with respect to the current utterance and context.
  • Perplexity Optimized Prompts
    • Given an LLM, we took the manually engineered prompt template, and created candidate prompt variants by using GPT3 and back translation.
    • We instantiated all such prompt templates using 100 instances (with full prompt sequence, including the input itself, and without the label)
    • Choose lowest perplexity template using the LLM.

Bishal Santra, Sakya Basak, Abhinandan De, Manish Gupta, Pawan Goyal. Frugal Prompting for Dialog Models. EMNLP Findings, 2023.

5

Manually engineered prompt template with summary of dialog history, persona and latest person1 utterance as dialog context and with one exemplar

6 of 13

Optimizing the Dialog History Input

  • Redundancies in conversations
    • back-channeling, clarification, and mistake correction.
    • responses from some dialog models (like textdavinci-003) could be elaborate and long.
  • Shortening Dialog Histories
    • Selection
      • Recent-k
      • Semantic-k: avg(SimCSE, Sentence Transformers)
    • Summarization
      • BART-D (DialogSum): facebook/bart-large 12L+12L
      • Pegasus-DS (DialogSum and SAMSum): google/pegasus-cnn_dailymail 16L+16L
      • Pegasus-CD (CNN/DailyMail): google/pegasus-cnn_dailymail 16L+16L
  • Shortening Background Information
    • BART, Pegasus.

Bishal Santra, Sakya Basak, Abhinandan De, Manish Gupta, Pawan Goyal. Frugal Prompting for Dialog Models. EMNLP Findings, 2023.

6

7 of 13

Datasets

  • Multisession Chat (MSC)
    • Multiple chat sessions whereby the speaking partners learn about each other’s interests and discuss the things they have learnt from past sessions.
    • Each user is asked to play a role (persona)
    • 16,299 context response pairs
    • Avg #utterances per conversation: 11.9

Bishal Santra, Sakya Basak, Abhinandan De, Manish Gupta, Pawan Goyal. Frugal Prompting for Dialog Models. EMNLP Findings, 2023.

  • Topical Chat (TC)
    • Each pair of users is assigned one or more topics along with some facts or knowledge about the topic, and the users are asked to have a conversation about the topic.
    • knowledge sections associated with the conversations.
    • 7,512 context response pairs
    • Avg #utterances per conversation: 20

7

8 of 13

Models; Prompt Design; Metrics

  • Models: GPT-3 (text-davinci-003), FLAN-T5, T0, Tk-Instruct.
  • Input prompt settings
    • Zero shot versus few shot (1 exemplar)
      • exemplar is chosen based on the immediately previous utterances if available, else it is randomly chosen from the dataset.
      • For conv=ABCDEFG and Recent-4
        • Target response=G, current utterance=F, recent-4 dialog history=BCDE
        • Exemplar: Target response=F, current utterance=E, recent-4 dialog history=ABCD
    • Manually designed versus perplexity optimized prompts
    • Usage of dialog history:
      • full history
      • summarized dialog history
      • Recent-k or Semantic-k
    • With and without summarized background-information

Bishal Santra, Sakya Basak, Abhinandan De, Manish Gupta, Pawan Goyal. Frugal Prompting for Dialog Models. EMNLP Findings, 2023.

  •  

8

9 of 13

Input Lengths

Comparison of average input length for various representations of dialog prompts across the two datasets for the few shot setting.

Bishal Santra, Sakya Basak, Abhinandan De, Manish Gupta, Pawan Goyal. Frugal Prompting for Dialog Models. EMNLP Findings, 2023.

  •  

9

10 of 13

Absolute perf analysis

Bishal Santra, Sakya Basak, Abhinandan De, Manish Gupta, Pawan Goyal. Frugal Prompting for Dialog Models. EMNLP Findings, 2023.

  • GPT-3 is best; Tk-Instruct is worst.
  • Zero shot > few shot
  • Perplexity optimized < Manually engg
  • Best results with full dialog history for TC in most cases for DEB and METEOR.
  • For MSC, even prompts with summarized history seem to do very well.
  • Semantic-k performs better than Recent-k.
  • Semantic-k peaks at k=4. Recent-k peaks at k=8 or 10.
  • Adding BI to Pegasus-DS helps boost DEB and METEOR but hurts BLEURT.

10

11 of 13

UID Results and Analysis

Bishal Santra, Sakya Basak, Abhinandan De, Manish Gupta, Pawan Goyal. Frugal Prompting for Dialog Models. EMNLP Findings, 2023.

  • Manually engineered prompts are better
  • Semantic-1 is the best. UID decreases as we increase k.
  • For both Recent-k and Semantic-k, UID reduces with increase in k.
  • Adding BI to Pegasus-DS does not help.
  • Pegasus-DS and BART-D perform better than Pegasus-CD.
  • Summaries>Full dialog history
  • Few-shot UID < zero-shot UID

11

12 of 13

 

  • For both MSC and TC, for DEB and METEOR, as “a” is increased
    • Summary-based dialog history variants tend to become better
    • Recent-k and Semantic-k variants tend to become less impressive.
  • For BLEURT, ranking is in favour of Semantic-1 or 2 and Recent-1 or 2 for all “a”

Bishal Santra, Sakya Basak, Abhinandan De, Manish Gupta, Pawan Goyal. Frugal Prompting for Dialog Models. EMNLP Findings, 2023.

  • Recent-1 and Semantic-1 are the recommended ways to summarize the context information if cost is an important factor to the user.
  • If cost is of less importance, then longer dialog summaries such as Pegasus-CD and Semantic-4 are recommended approaches for MSC.

12

MSC

13 of 13

Conclusion

  • Explored the tradeoff between model performance and cost for dialog systems.
  • Optimal representation of dialog history is one that provides the highest amount of usable information per token.
  • Insights
    • Summaries > full history.
    • Recent-k or Semantic-k > summaries.
    • Semantic-1 is best from both accuracy as well as UID perspective.
    • Zero-shot > Few-shot.

13

Bishal Santra, Sakya Basak, Abhinandan De, Manish Gupta, Pawan Goyal. Frugal Prompting for Dialog Models. EMNLP Findings, 2023.