1 of 133

Conversational Recommender SystemsCataldo Musto, Ph.D.�University of Bari – Italy��

2 of 133

Ciao!

I am Cataldo Musto

Tenure-track Assistant Professor at the Department of Computer Science, University of Bari

Research Lines: Recommender Systems, Knowledge Graphs, Natural Language Processing

Twitter: @cataldomusto

2

3 of 133

Overview

  1. Background and Motivations
  2. Basics of Conversational Recommender Systems (CRS)
  3. General Architecture of CRSs
    • Input Processing - Natural Language Understanding
    • User Modeling
    • Recommendation
    • Output Processing - Natural Language Generation
  4. Evaluation of CRSs
  5. Open Points and Discussion

3

4 of 133

1.

Background and Motivations

Why is it important to talk about CRSs at RecSys Summer School?

4

5 of 133

Conversational Agents are not new

5

(*) Weizenbaum, Joseph. "ELIZA—a computer program for the study of natural language communication between man and machine." Communications of the ACM 9.1 (1966): 36-45.

1966: Eliza(*)

90s-2000s: Microsoft Clippy

6 of 133

Conversational Agents today

6

7 of 133

7

Mastercard Digital Assistant

ChatGPT

8 of 133

Why?

  • Practical Motivations

Need for more ‘natural’ interaction strategies

Need for interaction strategies that are more suitable in particular scenarios (i.e., while driving)

Need to make automatic some tasks (i.e., CRM)

  • Technological/Methodological Motivations

Better algorithms to process voice and audio

Better algorithms to understand user input

Better algorithms to handle natural language (both input and output)

8

9 of 133

Why?

  • Practical Motivations
    • Need for more ‘natural’ interaction strategies in AI applications
    • Need for interaction strategies that better adapt to particular scenarios (i.e., while driving)
    • Interest from companies, to make automatic some tasks
  • Technological/Methodological Motivations

9

10 of 133

Why?

  • Practical Motivations
    • Need for more ‘natural’ interaction strategies in AI applications
    • Need for interaction strategies that better adapt to particular scenarios (i.e., while driving)
    • Interest from companies, to make automatic some tasks
  • Technological/Methodological Motivations
    • Better algorithms to process voice and audio
    • Better algorithms to understand user input (i.e., intents)
    • Better algorithms to handle natural language (both input and output)

10

11 of 133

Basic Problem Formulation

  • Conversational Agent: Input
    • Dialogue history: last n utterances

(Optional) background knowledge

  • Conversational Agent: Output
    • Next utterance to interact with user (in each turn)

(Optional) Some action to be performed

recommending item, switching on the light, playing some music, etc.

11

12 of 133

Basic Problem Formulation

  • Conversational Agent: Input
    • Dialogue history: last n utterances
    • (Optional) background knowledge (about users, items….)
  • Conversational Agent: Output
    • Next utterance to interact with user (in each turn)
    • (Optional) Some action to be performed
      • recommending item, switching on the light, playing some music, etc.

12

13 of 133

Taxonomies of Conversational Agents

13

Conversational Agents

Objective

Open-domain

Goal-oriented

Architecture

Modular

End-to-End

Interaction

System-initiative

User-iniative

Mixed

14 of 133

Taxonomy of CAs : Objective

  • Open-Domain
    • Able to maintain generic, chit-chat conversations
    • Must be able to support a wide variety of topics

14

15 of 133

Taxonomy of CAs : Objective

  • Open-Domain
    • Able to maintain generic, chit-chat conversations
    • Must be able to support a wide variety of topics
  • Goal-Oriented
    • Made for a specific domain in mind
    • Must be able to guide the conversation to complete the user’s tasks
      • e.g., booking a flight ticket, recommending movies, etc.
      • Conversational Recommender Systems

15

16 of 133

Taxonomy of CAs : Architecture

  • Modular architecture
    • Made up of a set of components
      • Each component has a specific responsibility
      • The components are usually connected to form a pipeline
  • End-to-End architecture
    • A single model handles the entire conversation
      • Given an input message, generate a text response
      • Typically based on deep learning
      • No need to develop special-purpose components

16

17 of 133

Taxonomy of CAs : Interaction

  • Interaction Modes
    • System Active – User Passive (SAUP)
      • The system leads the conversation by asking questions to user
      • User can only respond to the questions directly
    • System Active – User Engages (SAUE)
      • The system asks questions and user responds to the questions
      • Both system and user also chit-chat.
    • System Active – User Active (SAUA)
      • Both system and user can lead the conversation and chit-chat
    • User Active – System Passive
      • User drives the conversation by asking questions to system

17

18 of 133

Taxonomy of CAs : Interaction

  • Interaction Modes
    • System Active – User Passive (SAUP)
      • The system leads the conversation by asking questions to user
      • User can only respond to the questions directly
    • System Active – User Engages (SAUE)
      • The system asks questions and user responds to the questions
      • Both system and user also chit-chat.
    • System Active – User Active (SAUA)
      • Both system and user can lead the conversation and chit-chat
    • User Active – System Passive
      • User drives the conversation by asking questions to system

18

19 of 133

Taxonomies of Conversational Agents

19

Conversational Agents

Objective

Open-domain

Goal-oriented

Architecture

Modular

End-to-End

Interaction

System-initiative

User-iniative

Mixed

20 of 133

Part 1- Take-home Messages

  • Conversational Agents (Cas) are not new
  • New wave of CAs is due to several factors
    • Technological Factors
    • Practical Factors
  • CAs may differ based on their goal, their architecture and the interaction strategy
    • Goal: open-domain vs goal oriented
    • Architecture: modular vs end-to-end
    • Interaction: based on who can take the initiative

20

21 of 133

2.

Basics of Conversational Recommender Systems

Basic Principles

21

22 of 133

A Conversational Recommender System is a software system that supports its users in achieving recommendation‐related goals through a multi‐turn dialogue(*).

22

(*) Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

23 of 133

A Conversational Recommender System is a software system that supports its users in achieving recommendation‐related goals through a multi‐turn dialogue(*).

23

Task Orientation

Multi-turn Interaction

(*) Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

24 of 133

Conversational Recommender Systems

24

25 of 133

CRSs – Research Trend

25

From «Tutorial on Conversational Recommendation Systems» – ACM RecSys 2021

26 of 133

CRSs – Research Trend

26

From «Tutorial on Conversational Recommendation Systems» – ACM RecSys 2021

First Work on Critiquing for Recommender Systems

27 of 133

2000s: first work on critiquing and CRSs

27

from: Burke, Robin. "Interactive critiquing for catalog navigation in e-commerce." Artificial Intelligence Review 18.3-4 (2002): 245-267.

Requirements

28 of 133

2000s: first work on critiquing and CRSs

28

Requirements

Critiquing

from: Burke, Robin. "Interactive critiquing for catalog navigation in e-commerce." Artificial Intelligence Review 18.3-4 (2002): 245-267.

29 of 133

2000s: first work on critiquing and CRSs

29

Refinement

Requirements

Critiquing

from: Burke, Robin. "Interactive critiquing for catalog navigation in e-commerce." Artificial Intelligence Review 18.3-4 (2002): 245-267.

30 of 133

2000s: first work on critiquing and CRSs

30

From: Thompson, Cynthia A., Mehmet H. Goker, and Pat Langley. "A personalized system for conversational recommendations." Journal of Artificial Intelligence Research 21 (2004): 393-428.

Dialogue starts with an initial query based on user model. Then, based on answers/feedbacks, new questions are generated, to relax and constrain other features and retrieve potential recommendations

31 of 133

CRSs – Research Trend

31

From «Tutorial on Conversational Recommendation Systems» – ACM RecSys 2021

Deep Learning Wave

32 of 133

from 2018: Deep Learning wave in CRSs

32

from: Zhang, Tong, et al. "Kecrs: Towards knowledge-enriched conversational recommendation system." arXiv preprint arXiv:2105.08261 (2021).

33 of 133

Recently: first commercial CRSs

33

34 of 133

Specific Problem Formulation

  • CRS: Input
    • Dialogue history: last n utterances
    • User Preferences
    • Knowledge about the Items and the Domain

  • CRS: Output
    • Next utterance to interact with user (in each turn)
    • (Eventually) Item Recommendation
    • (Optional) Explanation

34

35 of 133

Part 2- Take-home Messages

  • CRSs are a specialization of general Conversational Agents
    • Specificities: goal-oriented and multi-turn dialogues

  • Deep Learning Wave also impacted CRSs
    • Better dialogues, better recommendations
    • First commercial systems
    • Recent trend: end-to-end play the main role

35

36 of 133

3.

General Architecture of CRSs

What are the main components of a conversational recommender system?

36

37 of 133

Taxonomies of CAs - Recap

37

Conversational Agents

Objective

Open-domain

Goal-oriented

Architecture

Modular

End-to-End

Interaction

System-initiative

User-initiative

Mixed

38 of 133

Modular CRSs vs End-to-End CRSs

  • Modular CRSs
    • Early strategy (but still exploited and effective)
    • Each component/requirement is mapped to a different module
      • It is possible to have different technological solutions
    • Dialogue State Manager is the core of the interaction

38

39 of 133

Modular CRSs vs End-to-End CRSs

  • Modular CRSs
    • Early strategy (but still exploited and effective)
    • Each component/requirement is mapped to a module
      • It is possible to have different technological solutions
  • End-to-End CRSs
    • More recent strategy, based on deep learning
      • Conversation seen as a sequence of messages
      • Solutions exploited for sequence data are typically suitable
        • Session-based Recommender, Sequential Recommendations
    • Just one architecture that handles all the different problems

39

40 of 133

Conceptual Architecture of a CRS

40

Source: Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

41 of 133

Input Processing Module

41

Source: Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

42 of 133

User Modeling Modules

42

Source: Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

43 of 133

Recommendation Module

43

Source: Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

44 of 133

Output Generation Module

44

Source: Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

45 of 133

End-to-End Systems

45

In end-to-end systems, all the components (or almost all the components) are mapped to a single deep architecture

46 of 133

Research Questions and Directions in CRSs

  1. Input Processing
    • What kind of interaction shall a CRSs support?
      • Natural Language? Buttons? Mixed?
    • What is the most suitable interaction strategy?
      • Voice? Text? Other forms (i.e., handwritten)
    • How do we process and understand user inputs?
      • Intent Recognition
  2. User Modeling
    • How to model user informative needs and preferences?
      • Entities? Objective features? Subjective features?

46

47 of 133

Research Questions and Directions in CRSs

  1. Recommendation
    • When do we end the preference elicitation phase and start the recommendation process?
      • Dialogue State Management
    • What kind of information is needed to provide effective recommendations?
      • Previous dialogs, background knowledge, item features?
  2. Output Generation
    • How to manage user feedbacks and to continue the dialogue?
    • How to return recommendations?
    • Are explanations useful?

47

48 of 133

3-a.

Input Processing

How can we process user inputs to extract needs and preferences?

48

49 of 133

Input Processing Module

  • Input
    • Generic user input (voice, text, handwritten, etc.)
  • Computational Tasks
    • Acquire and decode the input from the user
    • Process the information contained in input messages
  • Output
    • Detection of the intent of the user
    • Update of the state of the dialog

49

50 of 133

User Input

  • Input of the user may arrive in different forms

50

Voice (Natural Language)

Text (Natural Language)

Text (Forms and Buttons)

51 of 133

User Input

  • Input of the user may arrive in different forms

51

Voice (Natural Language)

Text (Natural Language)

Text (Forms and Buttons)

Requires Speech-to-Text

Easier strategy

Requires Natural Language Understanding

52 of 133

User Input

  • Input of the user may arrive in different forms

52

Voice (Natural Language)

Text (Natural Language)

Text (Forms and Buttons)

Requires Speech-to-Text

Easier strategy

OUR FOCUS

53 of 133

Natural Language Understanding

53

From: Martina, Alessandro Francesco Maria, et al. "A Virtual Assistant for the Movie Domain Exploiting Natural Language Preference Elicitation Strategies." Adjunct Proceedings of the 30th ACM Conference on User Modeling, Adaptation and Personalization. 2022.

54 of 133

Natural Language Understanding

54

Let’s analyze the modules we need for NLU

55 of 133

NLU: Intent Recognition

  • First Task carried out by a CRSs
    • The informative need expressed in each utterance provided by the user shall be identified
  • Input: user utterance, n possible intents
  • Output: intent expressed by the utterance

55

What is the intent?

56 of 133

NLU: Intent Recognition

56

A list of possible intents

57 of 133

NLU: Intent Recognition

57

58 of 133

NLU: Intent Recognition

58

Intent: to provide preferences

59 of 133

NLU: Intent Recognition

59

Intent: to ask for a recommendation

60 of 133

NLU: Intent Recognition

60

Intent: to provide feedbacks

61 of 133

NLU: Intent Recognition

  • How to implement it?
    • Modular CRSs: explicit intent recognition
    • A dedicated algorithm takes care of recognizing the intent
      • Typically carried out as a text classification task

61

62 of 133

NLU: Intent Recognition

  • How to implement it?
    • Modular CRSs: explicit intent recognition
    • A dedicated algorithm takes care of recognizing the intent
      • Typically carried out as a text classification task
        • Input
          • A list of intents and a set of training examples for each intent
            • Different ways to express the same informative need
        • Output:
          • Classification of the most suitable intent
      • An explicit change in the dialogue state is triggered

62

63 of 133

NLU: Intent Recognition

  • Many out-of-the-box solutions are also available
    • Google Dialogflow
    • Facebook’s Wit.ai
    • IBM Watson

  • End-to-End CRSs: shallow intent recognition
    • An implicit change of the state of the dialogue is obtained (i.e., weights in the neural network change)

63

64 of 133

NLU: Entity Recognition

  • In goal-oriented systems (such as CRSs) some information is typically needed to complete the task
    • In our case, user preferences are needed to provide users with recommendations

64

65 of 133

NLU: Entity Recognition

  • In goal-oriented systems (such as CRSs) some information is typically needed to complete the task
    • In our case, user preferences are needed to provide users with recommendations
  • It is very common to express preferences and needs in the form of entities
    • Entities = names, persons, locations, also key characteristics
    • Entity Recognition is the process which is carried out to identify entities in a text

65

66 of 133

NLU: Entity Recognition

66

Preferences are typically expressed in the form of entities

67 of 133

NLU: Entity Recognition

  • Many methods available
    • Dictionary
      • Pros: already available – Cons: maintaining and updating effort
    • Rules and Hand-written patterns (i.e., Telephone numbers, Addresses, etc.)
      • Pros: no training, no labeling - Cons: manual engineering, poor flexibility
    • Text Classification strategies (i.e., Random Forests, Naïve Bayes)
      • Pros: more flexible – Cons: requires labeled data and training
    • Sequence Models (i.e., Hidden Markov Models) and Deep Learning

  • Models and APIs available!

67

68 of 133

NLU: Sentiment Analysis

  • Given a message it is necessary to understand
    • The informative need of the user (intent recognition)
    • The informative elements that are mentioned (entity recognition)
    • The tone of the message

68

69 of 133

NLU: Sentiment Analysis

  • Given a message it is necessary to understand
    • The informative need of the user (intent recognition)
    • The informative elements that are mentioned (entity recognition)
    • The tone of the message

  • Sentiment Analysis algorithms carry out this task
    • Many solutions available
    • Commercial APIs, deep learning models

69

70 of 133

NLU: Sentiment Analysis

70

Preferences are typically expressed in the form of entities

71 of 133

NLU: Sentiment Analysis

71

Preferences are typically expressed in the form of entities

72 of 133

Input Processing: Take-home Messages

  • Input processing is the fundamental component of a CRS
  • The complexity of the implementation depends on how rich the dialogue shall be
    • Interaction based on buttons and form: relatively simple
      • Answers are pre-defined and bounded to a limited set of options
    • Natural Language Interaction: more sophisticated
      • Natural Language Understanding capabilities require solutions for intent recognition, entity recognition and sentiment analysis
      • In case of voice-based interaction, a speech-to-text solution is also required

  • The information collected in this step is used to continue the dialogue and trigger the recommendation process
    • Entities are exploited to model preferences and Intent is exploited to update the state of the dialogue

72

73 of 133

3-b.

User Modeling

How can we model user preferences and needs?

73

74 of 133

Intent Recognition - Recap

74

Some intents trigger the User Modeling module

75 of 133

User Modeling Module

  • Problem Formulation
    • Input
      • Content of user’s input messages
      • Entities and Sentiment extracted from the message
      • (eventually) state of the dialogue
    • Computational Tasks
      • Extract preferences and needs from input
    • Output
      • Representation of the profile of the user
      • (eventually) Update of the state of the dialog

75

76 of 133

User Modeling

76

Natural Language Understanding plays a key role, again

Objective Features

77 of 133

User Modeling

77

Natural Language Understanding plays a key role, again

Mention to items the user likes

78 of 133

User Modeling

78

Natural Language Understanding plays a key role, again

Mention to entities the

user likes

79 of 133

User Modeling

79

Natural Language Understanding plays a key role, again

Mention to entities the

user dislikes

80 of 133

User Modeling

80

Natural Language Understanding plays a key role, again

Also subjective features

81 of 133

Workflow for User Modeling in CRSs

81

A pipeline to extract both objective and subjective features from dialogue

From: Martina, Alessandro Francesco Maria, et al. "A Virtual Assistant for the Movie Domain Exploiting Natural Language Preference Elicitation Strategies." Adjunct Proceedings of the 30th ACM Conference on User Modeling, Adaptation and Personalization. 2022.

82 of 133

Workflow for User Modeling in CRSs

82

Knowledge Extraction is carried out in background. The resulting KB is then used in the dialogue.

From: Martina, Alessandro Francesco Maria, et al. "A Virtual Assistant for the Movie Domain Exploiting Natural Language Preference Elicitation Strategies." Adjunct Proceedings of the 30th ACM Conference on User Modeling, Adaptation and Personalization. 2022.

83 of 133

Knowledge Extraction

  • Objective properties
    • Non-controversial characteristics of the items
  • Typically collected by exploiting exogenous knowledge sources (i.e., knowledge graphs)
    • Based on Entity Linking techniques
      • Each entity is mapped to an entry in a knowledge graph
    • Based on the entities previously identified in text, entity linking is used to populate the profile of the user

83

84 of 133

DBpedia

84

Wikipedia

Unstructured Content

DBpedia

Structured Data

85 of 133

DBpedia and RDF

85

Starting from a mention to an entity from a dialogue, it is possible to populate the model of the user by also including properties connected to the entity

Richer representation!

86 of 133

Knowledge Extraction

  • Subjective properties regard users’ personal feelings when they enjoy the item
    • «Emotional Soundtrack», «Great acting»

86

87 of 133

Knowledge Extraction

  • Subjective properties regard users’ personal feelings when they enjoy the item
    • «Emotional Soundtrack», «Great acting»

  • Based on Aspect Extraction techniques
    • Extraction of key aspects from text (reviews, typically)
    • Basic solution
      • POS-tagging + Extraction of names OR Names + adjectives

87

88 of 133

User Modeling: Take-home Messages

  • Preferences of the user may have different forms
    • Objective properties
      • Non-controversial features
      • Available in exogenous knowledge sources (i.e., Dbpedia, Wikipedia)
    • Subjective properties
      • Personal feelings about the characteristics of the item (i.e., emotional soundtrack)
      • Opinion mining and aspect extraction techniques are needed
      • Useful to improve the quality of the preference elicitation process (*)

88

89 of 133

User Modeling: Take-home Messages

  • Modular CRSs
    • User Modeling output is passed to the recommendation model, as a list of features or as key-value pairs
    • User Modeling and recommendation are almost independent
    • User Modeling = preference elicitation

  • End-to-End CRSs,
    • Based on sequence data
    • User Model = sequence of concepts, entities, tokens the user is interested in
    • Task: predict the next token (i.e., item) based on the input sequence

89

90 of 133

3-c.

Recommendation

How (and when) do we provide users with recommendations?

90

91 of 133

Intent Recognition - Recap

91

Some intents trigger the Recommendation module

92 of 133

Recommendations through Dialogue

92

93 of 133

Recommendations through Dialogue

93

User asks for a recommendation

System returns a recommendation

User provides a feedback (mixed interaction)

94 of 133

Recommendations through Dialogue

94

When

to recommend

What

to recommend

95 of 133

Recommendations: when

  • When to recommend?
    • The answer is not trivial. Influenced by two factors.
      • Interaction strategy (user passive / user active)
      • The flexibility of the dialogue (fixed /flexible)

95

Recap

96 of 133

Recommendations: when

  • User passive – Fixed dialogue
    • Slot-filling strategy

    • Very common in early attempts (*)
      • Fixed set of questions and order
      • Through each question, different information are acquired
      • Effective, but poorly flexible

96

(*) Jannach, Dietmar, and Gerold Kreutler. "Rapid development of knowledge-based conversational recommender applications with advisor suite." Journal of Web Engineering (2007): 165-192.

97 of 133

Recommendations: when

  • User active – Flexible Dialogue
    • The user starts the conversation
    • The system asks questions
    • The user answers
    • When the system feels confident, a recommendation is provided

  • Some ‘intelligence’ is used to
    • Decide what to ask next
    • Decide when to recommend
      • i.e., Based on a confidence score

97

98 of 133

Recommendations: when

  • User active – Flexible Dialogue
    • The user starts the conversation
    • The system asks questions
    • The user answers
    • When the system feels confident, a recommendation is provided

  • Some ‘intelligence’ is used to
    • Decide what to ask next
    • Decide when to recommend
      • i.e., Based on a confidence score

98

99 of 133

Recommendations: when

  • User active – Flexible Dialogue
    • The user starts the conversation
    • The system asks questions
    • The user can answer, but he can also chit-chat by providing new information

  • Some ‘intelligence’ is used to
    • Decide what to ask next
    • Decide when to recommend
      • i.e., Based on a knowledge graph

99

From: Moon, Seungwhan, et al. "Opendialkg: Explainable conversational reasoning with attention-based walks over knowledge graphs." ACL. 2019.

100 of 133

Recommendations: what

  • Different approaches based on the overall architecture

  • Modular CRSs
    • Recommendation algorithm is just a module
    • Every state-of-the art implementations is suitable
      • Collaborative algorithms not good (sparsity issues)
      • Knowledge-aware or pure Knowledge-based are more effective

100

101 of 133

Recommendations: what

  • Modular CRSs
  • Knowledge-based
    • Constraints indicated by the user are used to filter out the items in the catalogue
    • Constraints relaxation strategies may be needed

101

From: Burke, R., Hammond, K., and Young, B.: The FindMe Approach to Assisted Browsing. IEEE Expert: Intelligent Systems and Their Applications 12(4):32‐40, 1997.

102 of 133

Modular CRSs

  • Knowledge-aware
    • Need some ‘knowledge’ about the items (i.e., item features)
      • Available in «standard» recommendation settings (i.e., MovieLens)
    • Information from input processing and user modeling is used to drive the recommendation process
      • Entities, items, categories, etc.

102

Source: Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

103 of 133

Modular CRSs

  • Knowledge-aware
    • Need some ‘knowledge’ about the items (i.e., item features)
    • Information acquired in the input processing and user modeling is used to drive the recommendation process
  • Martina et al. (*) exploit Personalized PageRank based on the properties and the items mentioned in the preference elicitation process

103

(*) Martina, Alessandro Francesco Maria, et al. "A Virtual Assistant for the Movie Domain Exploiting Natural Language Preference Elicitation Strategies." Adjunct Proceedings of the 30th ACM Conference on User Modeling, Adaptation and Personalization. 2022.

104 of 133

Modular CRSs

  • Knowledge-aware
    • Need some ‘knowledge’ about the items (i.e., item features)
    • Information acquired in the input processing and user modeling is used to drive the recommendation process
  • Martina et al. (*) exploit Personalized PageRank based on the properties and the items mentioned in the preference elicitation process

104

(*) Martina, Alessandro Francesco Maria, et al. "A Virtual Assistant for the Movie Domain Exploiting Natural Language Preference Elicitation Strategies." Adjunct Proceedings of the 30th ACM Conference on User Modeling, Adaptation and Personalization. 2022.

105 of 133

End-to-End CRSs

  • Datasets for end-to-end CRSs increased in the last few years
    • Available for many domains
  • Used to train neural architectures
    • Task: to learn a dialogue model
      • given a sequence of tokens (dialogue), predict the next token (item ID)
    • Methods for sequence-based recommendations and session-based recommendations are suitable

105

Source: Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

106 of 133

End-to-End CRSs

  • How to transform a dialogue in a sequence of elements?
  • Many directions
    • All tokens?
    • Sequences of genres?
    • Sequences of entities and genres?
    • Is it useful to exploit sentiment analysis?
    • Pre-train embeddings with exogenous knowledge (i.e., Word2Vec, DBpedia)?

106

Zou, Jie, et al. "Improving conversational recommender systems via transformer-based sequential modelling." Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2022.

107 of 133

Recommendation: Take-home Messages

  • To generate recommendations is the core task in a CRS
  • Two main research questions
    • When to recommend
      • Many strategies developed
        • To identify ‘good’ questions
        • To decide when enough preferences have been collected
    • What to recommend
      • In Modular CRSs, every recommendation algorithm can be exploited
        • knowledge—aware are the most suitable
      • In End-to-end CRSs, a dialogue model is learnt. Recommendation is a part of the dialogue
        • Managed as a «next-token» prediction taken
        • Methods based on transformers are state-of-the-art here

107

108 of 133

3-d.

Output Generation

How and when interact with the user? How to return the recommendations?

108

109 of 133

Intent Recognition - Recap

109

Every intent requires an adequate answer from the system

110 of 133

Output Generation

  • Every utterance of the user (mapped to an intent) requires an adequate answer
  • The answers to some intents are quite trivial
    • ’Acquire preferences’, ‘revise preferences’ just require a feedback of the CAs. «Ok, I got your preference»

110

111 of 133

Output Generation

  • Every utterance of the user (mapped to an intent) requires an adequate answer
  • The answers to some intents are quite trivial
    • ’Acquire preferences’, ‘revise preferences’ just require a feedback of the CAs. «Ok, I got your preference»

  • The generation of other answers is more challenging
    • How to manage chit-chat?
      • Utterances unrelated to the previous question and/or to the recommendation goal
    • How to generate an explanation?

111

112 of 133

Output Generation

  • The architecture of the CRSs also influences the output generation
  • Modular CRSs
    • Easy answers, typically based on templates
    • Natural Language Generation = filling in the template with the information
      • «I suggest you to watch <item-id>»

112

113 of 133

Output Generation

  • The architecture of the CRSs also influences the output generation
  • Modular CRSs
    • Easy answers, typically based on templates
    • Natural Language Generation = filling in the template with the information
      • «I suggest you to watch <item-id>»
  • End-to-End CRSs
    • Generation of the answers depends on training data
    • Natural Language Generation = dynamic generation of a sequence of suitable tokens based on the information seen in the training
      • More flexible, more diversified, more complex to be implemented

113

114 of 133

Output Generation

  • Response Generation based on dialogue models today dominates the research
  • Li et al (*) use an encoder/decoder network to generate suitable answers

  • Research Question
    • Is it effective?

114

(*) Li, Raymond, et al. "Towards deep conversational recommendations." Advances in neural information processing systems 31 (2018).

115 of 133

Output Generation

  • Is it effective?
    • Jannach et al. (*) showed that manual annotators evaluated many sentences as not meaningful

115

(*): Manzoor, A., Jannach, D.: Conversational Recommendation based on End‐to‐end Learning: How Far Are We? Computers and Human Behavior Reports, 2021

116 of 133

Output Generation

  • Retrieval-based Output Generation
    • Investigated by Manzoor et al. (*)
    • Based on the intuition that similarity measures can be exploited to identify suitable answers from a corpus of recorded dialogues
      • Similar question = Similar answer
      • Retrieval of semantically similar answers

    • Interesting (but preliminary) experimental results

116

(*) A. Manzoor and D. Jannach. Generation‐based vs. Retrieval‐based Conversational Recommendation: A User‐Centric Comparison. In RecSys ’21, 2021

117 of 133

Output Generation

  • Explanation
    • User may ask for a clarification (‘Why is it suitable for me? ’)
    • Modular CRSs
      • Generate when «explanation» intent is detected
        • Implemented as a black-box explanation algorithm or as part of the recommendation algorithm (explainable RS)
    • End-to-End CRSs
      • Explanation is part of the dialogue model
        • The models learns to generate an explanation when some utterance related to explanation need is collected
        • Knowledge-aware models also exploit some exogenous knowledge to generate explanations

117

118 of 133

Output Generation

  • Explanation in Modular CRSs
    • In explainable recommendation algorithms, the models jointly learn to explain and recommend
    • Chen et al. (*) exploit review data to generate a suitable recommendation and a justification based on Recurrent Neural Networks (GRU)
    • Dialogue combines predefined templates and dynamic parts

118

(*) Chen, Zhongxia, et al. "Towards explainable conversational recommendation.Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence. 2021.

119 of 133

Output Generation

  • Explanation in Modular CRSs
    • Martina et al. (*) generate explanations based on ExpLOD (^)
    • Overlapping descriptive properties are used for explanation
    • Completely independent from the recommendation algorithm

119

(*) Martina, Alessandro Francesco Maria, et al. "A Virtual Assistant for the Movie Domain Exploiting Natural Language Preference Elicitation Strategies." Adjunct Proceedings of the 30th ACM Conference on User Modeling, Adaptation and Personalization. 2022.

(^) Musto, Cataldo, et al. "Generating post hoc review-based natural language justifications for recommender systems." User Modeling and User-Adapted Interaction 31 (2021): 629-673.

120 of 133

Output Generation: Take-home Messages

  • Output generation can be generated in different way
  • Template-based Generation
    • Easy to implement and effective
    • Very static. Can the dialogue boring and repetitive.
  • Dynamic Generation based on Deep Learning
    • Based on corpora of recorded dialogs
      • A model of dialogue is learnt
    • More dynamic, but answers can be also poorly significant
    • New direction: retrieval-based answer generation
  • Explanations are fundamental as well
    • May be independent or encoded in the recommendation/dialogue

120

121 of 133

4.

Evaluation of Conversational RSs

How can we evaluate a Conversational Recommender System?

121

122 of 133

Evaluation of CRSs

  • Evaluation of CRSs is not trivial
  • What makes a CRS ‘good’ ?
    1. Accurate recommendations, obtained with less interaction turns?
    2. Engaging interaction, mimicking human-human dynamics?
    3. Realistic dialogue generation?

  • Still an open research problem

122

123 of 133

Evaluation of CRSs

  • Evaluation of CRSs is not trivial
  • Usual Dichotomy
    • Research Goals
      • High prediction accuracy, high effectiveness of the algorithms, ability at mimicking human dialogues
    • Business Goals
      • More sales? More profit? More engagement?

  • Still an open research problem

123

124 of 133

Evaluation of CRSs

  • Current evaluation paradigms
    • On-line Studies
      • Experimental research with users (i.e., A/B tests)
    • Off-line Studies
      • Evaluation of the accuracy of the single components
      • Simulation of the behavior of the users
    • Qualitative Studies

124

125 of 133

Evaluation of CRSs

  • Evaluation dimensions (*)
    • Effectiveness
      • Can we get good/better recommendations with using a conversational interaction?
      • Metrics: task completion rate, hit rate, choice satisfaction (ResQue Questionnaire), accuracy metrics (Precision, NDCG…)
    • Efficiency
      • Does a CRSs allow to be fasted in getting a good recommendation?
      • Metrics: interaction time, perceived effort, number of interactions in a conversation, cognitive effort, etc.

125

(*): Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

Metrics can be both qualitative and quantitative

126 of 133

Evaluation of CRSs

  • Evaluation dimensions (*)
    • Effectiveness
      • Can we get good/better recommendations with using a conversational interaction?
      • Metrics: task completion rate, hit rate, choice satisfaction (ResQUE questionnaire), accuracy metrics (precision, NDCG…)
    • Efficiency
      • Does a CRSs allow to be fasted in getting a good recommendation?
      • Metrics: interaction time, perceived effort, number of interactions in a conversation, cognitive effort, etc.

126

(*): Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

Metrics can be both qualitative and quantitative

127 of 133

Evaluation of CRSs

  • Evaluation dimensions (*)
    • Quality of Conversation
      • Is the conversation natural and fluent? Quality and relevance of sentences
      • Metrics: fluency, perplexity, BLEU Score, ROUGE score, SUS usability score, misrecognition rate, quality of dialogue, user control, etc.

127

(*): Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

Metrics can be both qualitative and quantitative

128 of 133

Evaluation of CRSs

  • Evaluation dimensions (*)
    • Quality of Conversation
      • Is the conversation natural and fluent? Quality and relevance of sentences
      • Metrics: fluency, perplexity, BLEU Score, ROUGE score, SUS usability score, misrecognition rate, quality of dialogue, user control, etc.

    • Effectiveness of Subtasks
      • Accuracy of the modules (i.e., intent recognizer, entity recognizer, explanation, number of critiques, etc.

128

(*): Jannach, Dietmar, et al. "A survey on conversational recommender systems." ACM Computing Surveys (CSUR) 54.5 (2021): 1-36.

Metrics can be both qualitative and quantitative

129 of 133

5.

Open Points and Discussion

Can we sketch open points and future research directions?

129

130 of 133

Conversational Recommender Systems

130

Search

Recommendation

Conversational Recommendation

User's Intention is clear, explicitly indicated by query

User's Intention is unclear, implicitly revealed in history 

User’s Intention is obtained through conversation

131 of 133

Open Points and Discussion

  • Research in the area of Conversational Recommender Systems is very active
    • Recent advances in NLP are the cornerstone of this trend
      • Business goals influence the research as well
      • Accurate and effective technologies have been released in the last few years (i.e., ChatGPT)

    • General dichotomy: modular vs end-to-end systems
      • Modular CRSs borrow problems and solutions from NLP e HCI
      • End-to-End Systems largely exploit deep learning and sequence data

131

132 of 133

Open Points and Discussion

  • Many open research points and research directions
  • Input Processing
    • What is the most effective strategy to acquire and handle user input?
  • User Modeling
    • What kind of information and features are worth to be encoded in a profile?
  • Recommendation
    • When to generate recommendation? How to handle long-term user interests?
    • Is it necessary to generate explanations?
  • Output Generation
    • Is it better to exploit static methods or strategies based on DL?
  • Evaluation
    • What makes a Conversational Recommender System useful and effective?

132

133 of 133

Thanks!

Any questions?

You can find me at:

@cataldomusto

cataldo.musto@uniba.it

133