Off-policy Learning over Heterogeneous Information for Recommendation
UTS CRICOS 00099F
Xiangmeng Wang 1, Qian Li 2, Dianer Yu 1, Guandong Xu 1
1 Data Science and Machine Intelligence (DSMI) Lab, School of Computer Science, University of Technology Sydney
2 School of Electrical Engineering, Computing and Mathematical Sciences, Curtin University
Presentation Outline
Background: Reinforcement learning (RL) based-Recommendation
Content Discovery
User dynamics
Long-term User Satisfaction
Recommender
Rewards (click?)
Action (rec items)
[1] Koren Y, Bell R, Volinsky C. Matrix factorization techniques for recommender systems[J]. Computer, 2009, 42(8): 30-37.
Learn target
policy
Target policy
Maximized
Reward
User
RS agent
Action
State
Recommender
Reward
User
Action
rare actions
actions
logging actions
Distribution drift
bias issue
Background: Off-policy Learning and the Bias Issue
User feedback collection
Policy deployment
Logging Policy
Target Policy
Presentation Outline
Motivation: Existing Bias Correction Works
Logged actions
...
...
Recommend actions
[2] Williamson E J, Forbes A. Introduction to propensity scores[J]. Respirology, 2014, 19(5): 625-635.
Motivation: Debiasing with Context Information
Logged actions
...
Agent
Heterogeneous Information Network
Recommend
action
Presentation Outline
Off-policy Learning over Heterogeneous Information for Recommendation (HINpolicy)
Main component
Methodology: Model Framework
Methodology: Context Representation Learning
Methodology: HIN-enhanced policy learning
Methodology: Counterfactual Risk Minimization
Presentation Outline
Evaluation: Recommendation performance
The SOTA model
Table. Performance comparison with baselines.
Evaluation: Ablation Study
Table. HINpolicy performance with (w/) or without (w/o) HIN
Presentation Outline
Conclusion
Questions?
UTS CRICOS 00099F
Get in touch:
Xiangmeng Wang (xiangmeng.wang@student.uts.edu.au)
Dr. Qian Li ( qli@curtin.edu.au)
Prof. Guandong Xu (Guandong.Xu@uts.edu.au)