1 of 48

PREDICTING TELECOM CHURNERS

A CASE OF MONGOLIA

May 2, 2018

Presenters:

Melody Sumiya

Shruti Bangad

2 of 48

CONTENTS

II. DATASET, VARIABLE DESCRIPTION, AND DESCRIPTIVE ANALYSIS

III. SELECTED MODELS

I. OBJECTIVE OF THE RESEARCH

V. LOGISTIC REGRESSION

VI. DECISION TREE

IV. CLUSTER ANALYSIS

VII. MODELS TESTING

VIII. CONCLUSION AND SUGGESTIONS

3 of 48

1

OBJECTIVE OF THE RESEARCH

Purpose and reasoning

4 of 48

Subscribers who discontinue their subscriptions to that service within a given time period

OBJECTIVE OF THE RESEARCH

PURPOSE AND REASONING

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

Definition

CHURNERS

# CHURNERS

# NEW CUSTOMER

Minimize

  • To categorize the customers in different groups based on their monthly usage.
  • To build predictive models to predict possible churners and compare the performance.
  • Apply in-class knowledge on practical case and get hands-on experience

PURPOSE OF THE RESEARCH

5 of 48

2

DATASET, VARIABLES DESCRIPTION AND DESCRIPTIVE ANALYSIS

Data manipulation

6 of 48

DATASET

VARIABLES DESCRIPTION

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

April 2013

TIME FRAME

VARIABLES

Contract ID – Customer ID (unique identifier)

Tenure – number of days using the service

Service type – A plan that customer has subscribed for (1-Classic, 0-Premium)

Location - Location of the customer (1-UB(capital city), 0-ON(countryside))

Cust type – whether the customer subscribed personally or via company(1-P, 0-C)

CDR count - number of days that any service was used in the month

Package (amt) - Basic plan amount (Minimum that you need to pay every month)

VAS (amt) - Additional amount that you need to pay for using Value added services

On-network, off-network (amt, cnt) - amount charged for calling in-network (on top of the base plan)

Data (amt, cnt) - amount charged for internet usage (on top of the base plan)

SMS on-network, off-network (amt, cnt) - amount charged for in-network sms (on top of the base plan)

Categorical

3

Numerical

15

Churn status (0-Churned or 1- Active)

Response variable

Total number of predictor variables

18

7 of 48

DATASET

DESCRIPTIVE ANALYSIS

By service type

By service ownership type

By location

Average usage of a subscriber by service type

Total number of subscribers

84,199

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

8 of 48

DATASET

DESCRIPTIVE ANALYSIS

Pearson correlation matrix

Chi-square test

Product

4.478***

Customer type

860.19***

Location

1167.7***

Pearson Chi-square test for correlation

*** 1% Significance, ** 5% Significance, * 10% Significance

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

9 of 48

Those who churned this month have NO DATA to use

DATASET

DATA CHALLENGE

This month’s status

Next month’s status

1

0

1

1

0

0

0

1

Churn status

Active – 1

Churned – 0

Use this month’s data to predict next month’s churn status

Next month’s Churn status

Those who were active in this month

84,199

Total number of subscribers

Those who were active

77,914

Churners vs. active subscribers

1.8%

2,682

Dataset

Random sampling to

balance the dataset

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

10 of 48

3

SELECTED MODELS

Suitable models

11 of 48

SELECTED MODELS

SUITABLE MODELS

Churn status

Active – 1

Churned – 0

Response variable

Logistic regression

Decision tree

Cluster analysis

Prediction

Exploration

Sample size

2,682

Test

805 (30%)

Training

1,877

(70%)

Sample size

2,682

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

12 of 48

4

CLUSTER ANALYSIS

13 of 48

Only for quantitative variables

Normalization – BBmisc

K-means - kmeans

Cluster number selection - mclus

CLUSTER ANALYSIS

MODEL SETUP

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

SOFTWARE

UNSTANDARDIZED

STANDARDIZED

Coded as dummy variables

CATEGORICAL

METHOD

K-means

14 of 48

CLUSTER ANALYSIS

NUMBER OF CLUSTERS

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

Within Groups Sum of Squares (WSS)

Within Groups Sum of Squares (WSS)

Difference between WSS

Difference between WSS

STANDARDIZED

UNSTANDARDIZED

K with lowest BIC

1

K with lowest BIC

2

3

6

3

2

6

2

3

6

8

4

8

4

8

Number of clusters

6

8

6

6

15 of 48

CLUSTER ANALYSIS

DATASET SELECTION

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

STANDARDIZED

UNSTANDARDIZED

Proportion of number of observations in each cluster

CHOSEN

16 of 48

CLUSTER ANALYSIS

PROFILING

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

CLUSTER 6:

Churners who live in capital city and use classic service

CLUSTER 5:

Half churners and active users who subscribed to premium service

CLUSTER 4:

Mostly active users who subscribed personally

CLUSTER 3:

Dominated by active users who subscribed to classic service

CLUSTER 2:

Active users who live in countryside and use classic service

CLUSTER 1:

Dominated by active users who subscribed to premium service

17 of 48

CLUSTER ANALYSIS

PROFILING

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

Number of days

% change from average

Number of active days

% change from average

Number of active days

% change from average

Number of active days

% change from average

18 of 48

CLUSTER ANALYSIS

PROFILING

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

IN NETWORK VOICE CALL

OFF NETWORK VOICE CALL

IN NETWORK SMS

OFF NETWORK SMS

MOBILE INTERNET

19 of 48

CLUSTER ANALYSIS

PROFILING

LOYALISTS

Cluster 2

NETIZENS

Cluster 4

TEXTERS

Cluster 1

CHATTERS

Cluster 3

DOUBTERS

Cluster 5

CHURNERS

Cluster 6

Mostly lives

in countryside

used the plan ~ 22 months

Subscribed to CLASSIC plan

  • Uses each service no more than average amount
  • Knows the value of the service
  • Chosen the right product
  • Pays 18,529₮ monthly

Active ~ 27

days per month

Lives in city

and countryside

used the plan ~ 32 months

Active ~ 26

days per month

Mostly subscribed to PREMIUM plan

  • Uses ~ 30 times greater mobile data
  • Pays 60,020₮ monthly, almost half goes to additional internet service
  • Uses little bit above average

Lives in city

and countryside

used the plan ~ 25 months

Active ~ 27

days per month

Mostly subscribed to PREMIUM plan

  • Sends text message 5 times greater
  • Uses greater value added services
  • Makes calls greater than average
  • Pays 44,406₮ monthly

Lives in city

and countryside

used the plan ~ 39 months

Active ~ 27

days per month

Subscribed to CLASSIC plan

  • Makes high amount of off network call (prob. International calls)
  • Uses each service little bit greater than average amount
  • Pays 33,937₮ monthly

Low Mid High

Probability to churn

Lives in city

and countryside

used the plan ~ 25 months

Active ~ 27

days per month

Subscribed to PREMIUM plan

  • Makes high amount of calls in and off network
  • Uses any other services at average
  • Pays 39,193₮ monthly

Lives in city

used the plan ~ 28 months

Active ~ 13

days per month

Mostly subscribed to CLASSIC plan

  • Uses any other services below average
  • Pays 13,436₮ monthly

CONCENTRATION

CONCENTRATION

CHEAP MOBILE DATA

SATISFIED

CHEAP INTERN. CALL

20 of 48

5

LOGISTIC REGRESSION RESULT

21 of 48

LOGISTIC REGRESSION

MODEL SETUP

Churn status (Churned or active)

Active – 1

Churned – 0

Response variable

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

Stepwise Logistic Regression for event = ‘0’

To find the probability of the user to churn in next month

SOFTWARE

OBJECTIVE

METHOD

22 of 48

LOGISTIC REGRESSION

RESULTS FOR STEPWISE LOGISTIC REGRESSION

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

 

 

23 of 48

LOGISTIC REGRESSION

RESULTS FOR STEPWISE LOGISTIC REGRESSION

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

GLOBAL TEST: LIKELIHOOD RATIO TEST

ODDS RATIO ESTIMATES

24 of 48

LOGISTIC REGRESSION

MODEL FIT

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

HOSMER AND LEMESHOW GOODNESS-OF-FIT TEST

GROUPS = 10

GROUPS = 5

MCFADDEN R2 INDEX 

25 of 48

LOGISTIC REGRESSION

ROC CURVE

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

AUC score: 0.894

26 of 48

LOGISTIC REGRESSION

MODEL TESTING

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

ACCURACY OF CLASSIFICATION = 0.8156956

CONFUSION MATRIX

Churn status (Churned or active)

Active – 0

Churned – 1

Response variable

Active

Churned

Active

348

54

Churned

101

302

  • SENSITIVITY = 74.94%
  • SPECIFICITY= 86.57%
  • FALSE POSITIVE RATE = 13.43%
  • FALSE NEGATIVE RATE = 26.06%

PREDICTED VALUE

ACTUAL VALUE

SOFTWARE

27 of 48

LOGISTIC REGRESSION

MODEL TESTING

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

ACCURACY OF CLASSIFICATION = 0.8347826

CONFUSION MATRIX WITH CUT OFF 0.4

  • SENSITIVITY = 83.13%
  • SPECIFICITY= 83.83%
  • FALSE POSITIVE RATE = 16.16%
  • FALSE NEGATIVE RATE = 16.87%
  • WE OBSERVE THAT NOW THE MODEL IS EQUALLY GOOD AT PREDICTING IF THE USER WILL CHURN OR NOT CHURN
  • 83.5 PERCENT OF THE TIMES THE MODEL WILL PREDICT ACCURATELY

Active

Churned

Active

337

65

Churned

68

335

PREDICTED VALUE

ACTUAL VALUE

28 of 48

6

DECISION TREE

29 of 48

DECISION TREE

ALGORITHM

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

Information Gain = Entropy of parent − Weighted average of entropy of child

HOW DOES THE DECISION TREE ALGORITHM DECIDE WHICH VARIABLE TO SPLIT FIRST?

Entropy measures the amount of information in a random variable

ENTROPY

 

30 of 48

Classification tree

DECISION TREE

MODEL SETUP

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

Tree - RPart

SOFTWARE

UNSTANDARDIZED

Coded as dummy variables

CATEGORICAL

METHOD

STEPS

Begin with a small cp

I

Pick the tree size that minimizes misclassification rate (i.e. prediction error)

II

Prune the tree using the best cp

III

Churn status

Active – 1

Churned – 0

Response variable

31 of 48

DECISION TREE

STEP 1: BEGIN WITH A SMALL CP

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

0.001

INITIAL CP

0.49973

ROOT NODE ERROR

32 of 48

DECISION TREE

STEP 1: BEGIN WITH A SMALL CP

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

33 of 48

DECISION TREE

STEP 2: PICK THE TREE SIZE

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

0.001

INITIAL CP

0.49973

ROOT NODE ERROR

BEST CP

34 of 48

DECISION TREE

STEP 3: PRUNE THE TREE

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

35 of 48

DECISION TREE

STEP 4: PREDICTION

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

Churned

Active

Churned

19

384

Active

49

353

PREDICTED VALUE

ACTUAL VALUE

CONFUSION MATRIX

Churned

Active

Churned

2%

48%

Active

6%

44%

PREDICTED VALUE

ACTUAL VALUE

CONFUSION MATRIX %

Accuracy

46%

  • SENSITIVITY = 47.9%
  • SPECIFICITY= 27.9%
  • FALSE POSITIVE RATE = 52.1%
  • FALSE NEGATIVE RATE = 72.1%

36 of 48

Regression tree

DECISION TREE

MODEL SETUP

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

Tree - Rpart

SOFTWARE

UNSTANDARDIZED

Coded as dummy variables

CATEGORICAL

METHOD

STEPS

Begin with a small cp

I

Pick the tree size that minimizes misclassification rate (i.e. prediction error)

II

Prune the tree using the best cp

III

Churn probability

1 – Prob. to be active

0 – Prob. to churn

Response variable

37 of 48

DECISION TREE

STEP 2: BEFORE PRUNING

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

0.001

INITIAL CP

0.25

ROOT NODE ERROR

38 of 48

DECISION TREE

STEP 2: BEFORE PRUNING

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

39 of 48

DECISION TREE

STEP 2: BEFORE PRUNING

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

0.001

INITIAL CP

0.25

ROOT NODE ERROR

BEST CP

40 of 48

DECISION TREE

STEP 3: PRUNE THE TREE

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

41 of 48

DECISION TREE

STEP 4: PREDICTION

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

Churned

Active

Churned

198

205

Active

55

347

PREDICTED VALUE

ACTUAL VALUE

CONFUSION MATRIX

Churned

Active

Churned

25%

25%

Active

7%

43%

PREDICTED VALUE

ACTUAL VALUE

CONFUSION MATRIX %

Accuracy

68%

  • SENSITIVITY = 61.9%
  • SPECIFICITY= 78.1%
  • FALSE POSITIVE RATE = 37.1%
  • FALSE NEGATIVE RATE = 21.7%

42 of 48

7

MODELS TESTING

43 of 48

MODELS TESTING

COMPARISON

LOGISTIC REGRESSION

Accuracy score: 83.4%

Misclassification error: 16.6%

AUC score: 0.894

CLASSIFICATION TREE

REGRESSION TREE

Accuracy score: 46.2%

Misclassification error: 53.8%

Accuracy score: 68%

Misclassification error: 32%

44 of 48

8

CONCLUSION

45 of 48

Prepping dataset is crucial

We had several month’s of information. The dataset was unbalanced. The model results were highly influenced by selection of month, time period, variable generation or coding etc.

CONCLUSION

GENERAL CONCLUSION

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

Logistic regression is the best out of three models

Logistic regression returned the highest accuracy among three models and better AUC score. Stepwise process helped us to select influential predictors, however, the goodness of fit was insignificant which leads to the next conclusion.

Dataset could be improved for this application

Low accuracy scores and the insignificant goodness of fit suggest that the dataset was not suitable for this application.

    • Number of period (months) we included was too small. We feel that for this kind of application, understanding past pattern is important.
    • Monthly accumulated data itself may not be appropriate. We argue that the data ought to be weekly or daily to track the behavior.

46 of 48

CONCLUSION

GENERAL CONCLUSION

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

We could add more important variables

We had 18 variables, 15 quantitative and 3 qualitative variables. Although we had enough variables to represent usage, we feel that some important variables were not accessible to us and weren’t included in the models. Such as:

    • Demographic variables (e.g. age, gender, marital status)
    • Service satisfaction attributes (e.g. satisfaction score, complaint, net promoter score)
    • Promotion attributes (e.g. type of promotion, duration, impact)
    • Detailed usage information(e.g. peak hour calls, sms, dropped call)

47 of 48

CONCLUSION

SHORTCOMINGS AND SUGGESTIONS

OBJECTIVE

DATASET

SELECTED MODELS

DECISION TREE

CLUSTER ANALYSIS

LOGISTIC REGRESSION

MODELS TESTING

CONCLUSION

Prepping dataset is crucial

Dataset could be improved for this application

We could add important variables

    • Include enough data and find way to work the data of more than one month
    • Modify the models to be able to pick up seasonality and other trends
    • Break down the data to weekly and compare with monthly models
    • Demographic variables (e.g. age, gender, marital status)
    • Service satisfaction attributes (e.g. satisfaction score, complaint, net promoter score)
    • Promotion attributes (e.g. type of promotion, duration, impact)
    • Detailed usage information(e.g. peak hour calls, sms, dropped call)

48 of 48

THANK YOU