1 of 27

LECTURE 2

dr. Jamolbek Mattiev

Application fields of data mining

2 of 27

Lecture outline

  • Introduction: data flood
  • Data mining application examples
  • Data mining & knowledge discovery
  • Data mining tasks

2 / 38

Data flood 🡪 application examples 🡪 terminology 🡪 data mining tasks 🡪 summary

3 of 27

Trends leading to data flood

  • More data is generated:
    • In business:�bank, telecom, other business transactions, …
    • In science:�astronomy, biology,�chemistry, …
    • On the web:�social networks,�e-commerce, …

3 / 38

●●●●●●● 🡪 application examples 🡪 terminology 🡪 data mining tasks 🡪 summary

4 of 27

Big data examples (15 years ago)

  • Europe's Very Long Baseline Interferometry (VLBI) has�16 telescopes, each of which produces 1 Gigabit/second�of astronomical data over a 25-day observation session
    • storage and analysis is a big problem;

  • AT&T handles billions of calls per day
    • so much data, it cannot be all stored – analysis has to be done “on the fly”, on streaming data;

4 / 38

●●●●●● 🡪 application examples 🡪 terminology 🡪 data mining tasks 🡪 summary

5 of 27

Largest databases in 2003

  • Commercial databases (Winter Corp. 2003 survey):
    • France Telecom has largest decision-support DB = ~30TB;
    • AT&T has database = ~26 TB;

  • Web:
    • Alexa internet archive: 7 years of data, 500 TB
    • Google searches 4+ Billion pages, many hundreds TB
    • IBM WebFountain, 160 TB
    • Internet Archive (www.archive.org), ~300 TB;

5 / 38

●●●●●●● 🡪 application examples 🡪 terminology 🡪 data mining tasks 🡪 summary

6 of 27

From terabytes to exabytes to …

  • UC Berkeley - estimate: 5 exabytes�(5 million terabytes) of new data was created in 2002.�www.sims.berkeley.edu/research/projects/how-much-info-2003/

  • US produces ~40% of new stored data worldwide.

  • 2006 estimate: 161 exabytes (IDC study)�www.usatoday.com/tech/news/2007-03-05-data_N.htm

  • 2010 projection: 988 exabytes.

6 / 38

●●●●●●● 🡪 application examples 🡪 terminology 🡪 data mining tasks 🡪 summary

7 of 27

Largest databases in 2005

Winter Corp. 2005 commercial DB survey:

    • Max Planck Inst. for Meteorology: 222 TB
    • Yahoo: ~100 TB (largest data warehouse)
    • AT&T: ~94 TB

http://dssresources.com/news/1010.php

7 / 38

●●●●●●● 🡪 application examples 🡪 terminology 🡪 data mining tasks 🡪 summary

8 of 27

Data growth

In 2 years,�the size of the largest database TRIPLED !!!

8 / 38

●●●●●●● 🡪 application examples 🡪 terminology 🡪 data mining tasks 🡪 summary

9 of 27

The present situation

  • Example: �end of June, 2017 – CERN‘s data center stores more than 200 petabytes of data�(200 million gigabytes)

9 / 38

●●●●●●● 🡪 application examples 🡪 terminology 🡪 data mining tasks 🡪 summary

10 of 27

Data growth rate

  • In the year 2002 – 2 times more data�has been produced than in the year 1999;

  • In the year 2005 – 3 times more data�has been produced than in the year 2003;

  • Very little data will ever be looked at by a human;

Knowledge Discovery and Data Mining�are NEEDED to make sense and use of data !!!

10 / 38

●●●●●●● 🡪 application examples 🡪 terminology 🡪 data mining tasks 🡪 summary

11 of 27

Lecture outline

  • Introduction: data flood
  • Data mining application examples
  • Data mining & knowledge discovery
  • Data mining tasks

11 / 38

Data flood 🡪 application examples 🡪 terminology 🡪 data mining tasks 🡪 summary

12 of 27

  • Science:
    • astronomy, bioinformatics, drug discovery. …
  • Business:
    • CRM (Customer Relationship management), fraud detection,�e-commerce, manufacturing, sports/entertainment, telecom, targeted marketing, health care, …
  • Web:
    • search engines, advertising, web and text mining, …
  • Government:
    • surveillance, crime detection, profiling tax cheaters, …

12 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

Machine learning/data mining �application areas

13 of 27

Application areas

What do you think are some of the most important and widespread business applications of data mining?

13 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

14 of 27

Data mining�for customer modeling

  • Attrition prediction,
  • targeted marketing (cross-sell, customer acquisition),
  • credit-risk,
  • fraud detection,
  • banking,
  • telecom,
  • retail sales, …

14 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

15 of 27

Customer attrition: case study

  • Situation:�Attrition rate for mobile phone customers is around 25-30% a year (US data)!

  • With this in mind, what is the DM task?
    • Assumption: we have customer information for the past N months.

15 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

16 of 27

Customer attrition: case study (2)

Task:

  • Predict who is likely to attrite next month.
  • Estimate customer value and what is the cost-effective offer to be made to this customer.

16 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

17 of 27

Customer attrition: results

  • Verizon Wireless built a customer data warehouse;
  • Identified potential attriters;
  • Developed multiple, regional models;
  • Targeted customers with high propensity to accept�the offer;
  • Reduced attrition rate from over 2%/month to under 1.5%/month (huge impact, with >30 M subscribers)

(Reported in 2003)

17 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

18 of 27

Assessing credit risk:�case study

  • Situation: Person applies for a loan.
  • Task: Should a bank approve the loan?

  • Note: �People who have the best credit don’t need the loans,�and people with worst credit are not likely to repay�🡺 bank’s best customers are “in the middle”.

18 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

19 of 27

Credit risk: results

  • Banks develop credit models using variety of machine learning methods,
  • mortgage and credit card proliferation are the results of being able to successfully predict if a person is likely to default on a loan,
  • widely deployed in many countries.

19 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

20 of 27

e-commerce

  • A person buys a book (product) at Amazon.com

What is the task?

20 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

21 of 27

Successful e-commerce:�case study

Task:�Recommend other books (products)�this person is likely to buy.

Amazon does clustering based on books bought:

    • Customers who bought “Advances in Knowledge Discovery and Data Mining”, also bought “Data Mining: Practical Machine Learning Tools and Techniques with Java Implementations”.
  • Recommendation program is quite successful.

21 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

22 of 27

Unsuccessful e-commerce:�case study (KDD-Cup 2000)

Data:

clickstream and purchase data from Gazelle.com, legwear and legcare e-tailer.

Question:

Characterize visitors who spend�more than $12 on an average order at the site.

Dataset = 3,465 purchases, 1,831 customers,

Very interesting analysis by Cup participants

    • thousands of hours – $X,000,000 (Millions) of consulting,

Total sales: -$Y,000,

Obituary: Gazelle.com out of business, Aug 2000.

22 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

23 of 27

Genomic microarrays:�case study

Given microarray data for a number of samples (patients), can we:

  • accurately diagnose the disease?
  • predict outcome for given treatment?
  • recommend best treatment?

23 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

24 of 27

Example: ALL/AML data

38 training cases, 34 test cases, ~7,000 genes

2 classes: Acute Lymphoblastic Leukemia (ALL) vs � Acute Myeloid Leukemia (AML)

Use train data to build diagnostic model

ALL

AML

Results on test data:

33/34 correct, 1 error (may be mislabeled)

24 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

25 of 27

Security and fraud detection: case study

  • Credit card fraud detection
  • Detection of money laundering
    • FAIS (US Treasury)
  • Securities fraud
    • NASDAQ KDD system
  • Phone fraud
    • AT&T, Bell Atlantic,�British Telecom/MCI
  • Bio-terrorism detection at Salt Lake Olympics 2002

25 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

26 of 27

Data mining and privacy

  • In 2006, NSA (National Security Agency) was reported�to be mining years of call info, to identify terrorism networks;
  • Social network analysis has a potential to find networks;
  • Invasion of privacy – do you mind if your call information is in a government database?
  • What if NSA program finds one real suspect for 1,000 false leads? 1,000,000 false leads?

26 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

27 of 27

  • require knowledge-based decisions
  • have a changing environment
  • have sub-optimal current methods
  • have accessible, sufficient, and relevant data
  • provides high payoff for the right decisions!

Privacy considerations are important�if personal data is involved !!!

27 / 38

Data flood 🡪 ●●●●●●●● 🡪 terminology 🡪 data mining tasks 🡪 summary

Problems�suitable for data mining