Predicting New York Times Best Seller using NLP
Samaye Lohan, Logan Heft, Dongjun Shin, & Youn Kyeong Chang
Motivation
Goal
Data(Web Scraped)
Discussion
[1] Toner Buzz: Eye-Popping Book and Reading Statistics. https://www.tonerbuzz.com/blog/book-and-reading-statistics/ [Online; accessed 11-Apr-2022] (2022)
[2] Statista: U.S. Book Industry/Market—Statistics & Facts. https://www.statista.com/chart/26572/average-number-of-books-read-by-us-residents-per-year/ [Online; accessed 11-Apr-2022] (2022)
[3] Xiaobing Sun and Wei Lu, “Understanding Attention for Text Classification “, Singapore University of Technology and Design(2020)
References
Results
3. Single Head Attention
4. Multi Head Attention
2. GRU + CNN
Model | Test accuracy |
GRU | 0.73 |
CNN+GRU | 0.81 |
Single-Head Attention | 0.83 |
Multi-Head Attention | 0.78 |
Model 1: Embedded textual input by incorporating two dense layers.
Model 2: GRU + convolutional layers to the textual input and applied max pooling to prevent overfitting
Model 3: CNN Layers were replaced with attention layers to increase performance in addition to implementing a bidirectional LSTM layer.
Model 4: Extended Model 3 to now include multi-headed attention to automatically figure out two-way relationships.
The publishing industry profits, like other cultural industries is highly dependent on extremely successful books. Publishers may find it helpful at the selection process of the publishing to predict the success of the book in advance to save money on production costs.
Input 1(Tokenized):Title
Input 2 (Scaled): Rating, # Of Pages,
Year Published
Input 3 (One Hot Encoded): Top Author
Label (One Hot Encoded): NYT Bestseller
Methodology
Create a model that predicts New York Times best sellers by title, top author, ratings, number of pages and the year published. Our goal accuracy was to achieve 85%
Our group wanted to achieve an accuracy of 85% as it was a realistic enough target given the vast nature of the meta information tied to books. Our initial approach was to use the “title” feature as the input in order to predict whether the book is a New York Times bestseller. However, we soon realized that our GRU model yielded a relatively low accuracy, 73%, due to the fact that we were only considering one textual input. To further improve our model, we concatenated multiple models together with different inputs, textual and integer, in addition to appending max pooling layers to reduce overfitting.