1 of 41

Crawling your own Dataset for Research using Python

Siddhartha Anand

2 of 41

Prerequisites

  • Love for
    1. Python
    2. Discovering through data

3 of 41

What will you know by the end of this talk?

You will know where to begin!

4 of 41

Contents

Roadmap

  • Introduction (3-4 min)
    • Why is the data required?
  • Already existing datasets (1-2 min)
    • Where is the data?
    • Limitations
  • What to do then? (2-3 min)
    • Crawling/Scraping
    • Ethical Issues
  • Static & Dynamic websites (0-1 min)
  • How to? - Tools (15-16 min)
    • Scrapy + bs4 (Dblp - SNA)
    • Scrapy + bs4 (stackoverflow - Text)
    • Scrapy + bs4 (Imdb - Text)
    • twitter api - SNA
    • facebook api - SNA
  • Q&A (3-4 min)

5 of 41

Introduction

Why is the data required?

  • Social Network Analysis
  • Sentiment Analysis
  • Weather Prediction
  • Financial forecasts
  • Recommendation Systems

6 of 41

Recommendation Systems (Answering WHY)

Why is the data required?

  • “Birds of a feather flock together.”
    • You might like what your friends like - food, books, sports, travel, movies, etc.
    • Analyze your social network - Recommend - Improve.
  • 80/20 Rule
    • You might be inclined to ‘like’ things which are similar to the things you have ‘liked’ before.
    • Cuisines, Authors, Places, Actors, Directors, Genres etc.
    • Analyze your history - Recommend - Improve.
  • Netflix/Amazon

7 of 41

8 of 41

Algorithmic trading: Financial Markets

Why is the data required?

  • In algorithmic trading, machine learning helps to make better trading decisions.
  • Machine learning algorithms can analyze thousands of data sources simultaneously, something that human traders cannot possibly achieve.
  • Given the vast volumes of trading operations, that small advantage often translates into significant profits.

9 of 41

Social Network Analysis

Why is the data required?

  • Product Advertising
  • Personalized Recommendations
  • Understanding complex and hidden human patterns
  • Finding and busting terrorist networks.

10 of 41

So, what’s the problem?

When do we NOT have a problem?

When the data is readily available to you.

  • If you are part of a high-profile government funded research project.

  • If you or your research team has collaborations with world class labs that provide such data readily to you.

11 of 41

So, what’s the problem?

When do we have a problem?

When the data is not available to you.

  • Usually, you do not have the required large-scale data to conduct your own independent research, (like me).
  • You would want to conduct some research in a completely diversified area but do not have the sanitised data for it (like me).
  • You want to compare state-of-the-art algorithms on new datasets being generated instead of referring to decade-old datasets (like me).

12 of 41

Where is Data?

Everywhere you look!�Structured/Unstructured form

  • Dblp - co-authorship network
  • Textual data
  • Twitter
  • Facebook
  • Zomato
  • Stackoverflow
  • Blogs
  • SDSS

13 of 41

Where is Data?

How to get it?

  • Some organizations/websites helpful.
    • API access (twitter)
    • Xml format (dblp)
    • Zip files (snap, kaggle)
    • Json format
    • FITS (Image Astronomical Data)

14 of 41

Static

&

Dynamic Websites

  • Static Websites
    • Easier to scrape/crawl.
    • No dynamic content helps in identifying the parts we need to crawl in a single HTTP request.
    • Examples

15 of 41

Static

&

Dynamic Websites

  • Dynamic Websites
    • Not so easy to scrape/crawl.
    • Needs to be monitored for parts which we require.
    • A more detailed analysis of the Request (using the Network tab in the browser).
    • Data that you require is dynamic in nature.
    • Examples

16 of 41

Ethical Issues

Issues to keep in mind before starting your scraper.

  • Go over robots.txt first and understand what can/cannot be crawled.
  • Limit the speed of your program to avoid DDOS attacks.
  • Identify yourself/your scraper and be friendly.

17 of 41

Examples

  • We will only be focusing on getting the data from different kinds of sources.
    • Twitter API
    • DBLP - static
    • Imdb Reviews - static
    • www.investing.com - dynamic

18 of 41

Twitter API

Using readily available APIs

19 of 41

Twitter API

Analysis of the sentiment of tweets tweeted by Modi.

  • Streaming API
    • Data sent in near-real time
    • PUSH from Twitter
    • Persistent Connection needs to be open/long-lived
  • REST API
    • PULL from Twitter
    • Every HTTP Request is a new one
    • Past data can be queried

20 of 41

Twitter for developers

21 of 41

Twitter for developers: Things to consider

  • Why is the app required?
    • An App is a gateway to the Twitter Data.
    • Almost all data access needed, has to be through registering an App.
    • An App registration provides you with keys that is used in every HTTP Request.
  • Authentication
    • CONSUMER_API_KEY
    • CONSUMER_API_SECRET
    • ACCESS_TOKEN
    • ACCESS_TOKEN_SECRET
  • Rate Limits
    • Exponential Backoff
    • Status Code 402
  • Usage
    • tweepy (python) to start accessing Twitter Data

22 of 41

Twitter for developers: Registering your App

23 of 41

Twitter for developers: Code

  • https://github.com/SiddharthaAnand/pycon2019/tree/master/dblp
    • Let’s try this out.

24 of 41

Co-authorship network

Crawling dblp

  • Go to /robots.txt
    • https://dblp.uni-trier.de/robots.txt
    • Understand what can/cannot be crawled.
  • Analyze the html content.
  • Use scrapy and bs4 to crawl.
  • Analyze the data!

25 of 41

Dblp - robots.txt

26 of 41

Dblp - Analyze the html

27 of 41

Dblp - Analyze the html

28 of 41

Dblp-Code

  • https://github.com/SiddharthaAnand/pycon2019/tree/master/dblp
    • Let’s try this out.
  • https://docs.scrapy.org/en/latest/
    • A fast and powerful scraper which can also be deployed to the cloud for large scale crawls.
  • https://www.crummy.com/software/BeautifulSoup/bs4/doc/
    • An amazing html/xml parser in python.

29 of 41

Imdb reviews

User reviews

Problem definition: You need textual user reviews to train your machine learning model which can classify new reviews as Positive or Negative.

Data: Get data from imdb!

Data that you require is static in nature.

30 of 41

Imdb reviews

User reviews

  • Go to /robots.txt
    • https://www.imdb.com/robots.txt
    • Understand what can/cannot be crawled.
  • Analyze the html content.
  • Use scrapy and bs4 to crawl.
  • Analyze the data!

31 of 41

imdb - robots.txt

32 of 41

imdb - Analyze the html

33 of 41

Imdb-Code

  • https://github.com/SiddharthaAnand/pycon2019/tree/master/dblp
    • Let’s try this out.
  • https://docs.scrapy.org/en/latest/
    • A fast and powerful scraper which can also be deployed to the cloud for large scale crawls.
  • https://www.crummy.com/software/BeautifulSoup/bs4/doc/
    • An amazing html/xml parser in python.

34 of 41

Dynamic Websites

www.investing.com

Financial Data

Problem: Predict the value of BTC for a future date using Machine Learning.

Data: Collect historical data from www.investing.com for training your model.

35 of 41

www.investing.com - Analyze the html

36 of 41

Analyzing Networks Tab

37 of 41

investing.com-Code

  • https://github.com/SiddharthaAnand/pycon2019/tree/master/investingdotcom
    • Let’s try this out.
  • https://www.crummy.com/software/BeautifulSoup/bs4/doc/
    • An amazing html/xml parser in python.
  • Analyzing the Networks tab gives us hints as to how to proceed further.
    • Replicate a Request same as the one by the browser.

38 of 41

Already available data-stores

Sometimes you might get exactly what you need.

  • snap.stanford.edu
  • Kaggle

39 of 41

40 of 41

kaggle

41 of 41

Q/A?