1 of 1

Predicting New York Times Best Seller using NLP

Samaye Lohan, Logan Heft, Dongjun Shin, & Youn Kyeong Chang

Motivation

Goal

Data(Web Scraped)

Discussion

[1] Toner Buzz: Eye-Popping Book and Reading Statistics. https://www.tonerbuzz.com/blog/book-and-reading-statistics/ [Online; accessed 11-Apr-2022] (2022)

[2] Statista: U.S. Book Industry/Market—Statistics & Facts. https://www.statista.com/chart/26572/average-number-of-books-read-by-us-residents-per-year/ [Online; accessed 11-Apr-2022] (2022)

[3] Xiaobing Sun and Wei Lu, “Understanding Attention for Text Classification “, Singapore University of Technology and Design(2020)

References

Results

3. Single Head Attention

  1. GRU

4. Multi Head Attention

2. GRU + CNN

Model

Test accuracy

GRU

0.73

CNN+GRU

0.81

Single-Head Attention

0.83

Multi-Head Attention

0.78

Model 1: Embedded textual input by incorporating two dense layers.

Model 2: GRU + convolutional layers to the textual input and applied max pooling to prevent overfitting

Model 3: CNN Layers were replaced with attention layers to increase performance in addition to implementing a bidirectional LSTM layer.

Model 4: Extended Model 3 to now include multi-headed attention to automatically figure out two-way relationships.

The publishing industry profits, like other cultural industries is highly dependent on extremely successful books. Publishers may find it helpful at the selection process of the publishing to predict the success of the book in advance to save money on production costs.

Input 1(Tokenized):Title

Input 2 (Scaled): Rating, # Of Pages,

Year Published

Input 3 (One Hot Encoded): Top Author

Label (One Hot Encoded): NYT Bestseller

Methodology

Create a model that predicts New York Times best sellers by title, top author, ratings, number of pages and the year published. Our goal accuracy was to achieve 85%

Our group wanted to achieve an accuracy of 85% as it was a realistic enough target given the vast nature of the meta information tied to books. Our initial approach was to use the “title” feature as the input in order to predict whether the book is a New York Times bestseller. However, we soon realized that our GRU model yielded a relatively low accuracy, 73%, due to the fact that we were only considering one textual input. To further improve our model, we concatenated multiple models together with different inputs, textual and integer, in addition to appending max pooling layers to reduce overfitting.