1 of 1

Mining Digital Media to Understand the

Individual and Societal Impact of Substance Use Disorders

Daniel Habib (Preceptors: Brenda Curtis, Salvatore Giorgi, Douglas Bellew, TingTing Liu)

Technology and Translational Research Unit, National Institute on Drug Abuse Intramural Research Program

Introduction

  • Stigma, privacy concerns, and legal repercussions drive many people with a substance use disorder (SUD) to seek support online (Bergman and Kelly, 2021).
  • Unlike other platforms such as Twitter, Reddit can be filtered by subforum (i.e., subreddit) and does not have a character limit, which proves more fruitful for language processing (Cohan et al., 2018).
  • Behavioral health datasets rarely include substance use (SU) data, and when they do, they only include the most common ones.

Methods

Aims

  • Create SU-related and self-disclosed SUD/recovery datasets of Reddit posts
  • Improve reproducibility and uniformity for assessing substance use via social media

Future Directions

Discussion

  • All subreddits and posts from 01/2006-12/2020 were extracted from the Pushshift Reddit Dataset and uploaded to MySQL (Baumgartner et al., 2020).
  • SU keywords were used to find SU subreddits and compile posts within them into the SU-related Dataset.
  • The Self-disclosed SUD and Recovery Dataset was compiled from Reddit users within the SU subreddits who self-disclosed an SUD diagnosis. All of these users’ posts both within and outside of SU subreddits were included in the dataset. Additionally, posts from users with similar traits within health and general subreddits were added as matched controls.

  • Compare how individuals with an SUD have been dehumanized on various media platforms
  • Optimize drug term discovery methods
  • Geolocate Reddit users to yield location-specific data (Balsamo et al., 2019)
  • Study the intersectionality of an SUD with gender or other SUDs
  • Generalize findings to different contexts (Huang et al., 2012)
  • Analyze users’ language to study/predict users with an SUD
  • SU-related Dataset: corpus for studying the language around SU
  • Self-disclosed SUD and Recovery Dataset: centralized source for testing algorithms that automatically detect/predict users with an SUD
  • These datasets fill a gap in behavioral health data that is not addressed by mental health datasets.

Figure 1. Schema of the SU-related Dataset and the Self-disclosed SUD and Recovery Dataset. Posts are within quotes and subreddits begin with “r/”. The SU-related Dataset is public facing while the Self-disclosed SUD and Recovery Dataset must requested.

Limitations

  • Real users with bot-like names
  • Capturing slang terms
  • False positives in diagnosis patterns
  • Matched controls
  • Representativeness of users who self-disclose

Ethical Considerations

  • Consent and anonymity
  • Responsibilities of social media companies