1 of 48

Data Mining �Association Analysis: FP-Growth algorithm

​

​

​

​

​

​

2 of 48

FP-growth Algorithm

  • Use a compressed representation of the database using an FP-tree

​

  • Once an FP-tree has been constructed, it uses a recursive divide-and-conquer approach to mine the frequent itemsets

3 of 48

FP-tree construction

null

A:1

B:1

null

A:1

B:1

B:1

C:1

D:1

After reading TID=1:

After reading TID=2:

4 of 48

FP-Tree Construction

null

A:7

B:5

B:3

C:3

D:1

C:1

D:1

C:3

D:1

D:1

E:1

E:1

Pointers are used to assist frequent itemset generation

D:1

E:1

Transaction Database

Header table

5 of 48

FP-growth

null

A:7

B:5

B:1

C:1

D:1

C:1

D:1

C:3

D:1

D:1

Conditional Pattern base for D: � P = {(A:1,B:1,C:1),� (A:1,B:1), � (A:1,C:1),� (A:1), � (B:1,C:1)}

Recursively apply FP-growth on P

Frequent Itemsets found (with sup > 1):� AD, BD, CD, ACD, BCD

D:1

6 of 48

Tree Projection

Set enumeration tree:

Possible Extension: E(A) = {B,C,D,E}

Possible Extension: E(ABC) = {D,E}

7 of 48

Tree Projection

  • Items are listed in lexicographic order
  • Each node P stores the following information:
    • Itemset for node P
    • List of possible lexicographic extensions of P: E(P)
    • Pointer to projected database of its ancestor node
    • Bitvector containing information about which transactions in the projected database contain the itemset

​

8 of 48

Projected Database

Original Database:

Projected Database for node A:

For each transaction T, projected transaction at node A is T ∩ E(A)

9 of 48

ECLAT

  • For each item, store a list of transaction ids (tids)

​

​

​

TID-list

10 of 48

ECLAT

  • Determine support of any k-itemset by intersecting tid-lists of two of its (k-1) subsets.

​

​

​

​

​

​

  • 3 traversal approaches:
    • top-down, bottom-up and hybrid
  • Advantage: very fast support counting
  • Disadvantage: intermediate tid-lists may become too large for memory

∧

→

11 of 48

Rule Generation

  • Given a frequent itemset L, find all non-empty subsets f ⊂ L such that f → L – f satisfies the minimum confidence requirement
    • If {A,B,C,D} is a frequent itemset, candidate rules:

ABC →D, ABD →C, ACD →B, BCD →A, �A →BCD, B →ACD, C →ABD, D →ABC�AB →CD, AC → BD, AD → BC, BC →AD, �BD →AC, CD →AB, �

  • If |L| = k, then there are 2k – 2 candidate association rules (ignoring L → ∅ and ∅ → L)

12 of 48

Rule Generation

  • How to efficiently generate rules from frequent itemsets?
    • In general, confidence does not have an anti-monotone property

c(ABC →D) can be larger or smaller than c(AB →D)

​

    • But confidence of rules generated from the same itemset has an anti-monotone property
    • e.g., L = {A,B,C,D}:� � c(ABC → D) ≥ c(AB → CD) ≥ c(A → BCD)

      • Confidence is anti-monotone w.r.t. number of items on the RHS of the rule

13 of 48

Rule Generation for Apriori Algorithm

Lattice of rules

Pruned Rules

Low Confidence Rule

14 of 48

Rule Generation for Apriori Algorithm

  • Candidate rule is generated by merging two rules that share the same prefix�in the rule consequent

​

  • join(CD=>AB,BD=>AC)�would produce the candidate�rule D => ABC

​

  • Prune rule D=>ABC if its�subset AD=>BC does not have�high confidence

15 of 48

Effect of Support Distribution

  • Many real data sets have skewed support distribution

Support distribution of a retail data set

16 of 48

Effect of Support Distribution

  • How to set the appropriate minsup threshold?
    • If minsup is set too high, we could miss itemsets involving interesting rare items (e.g., expensive products)

​

    • If minsup is set too low, it is computationally expensive and the number of itemsets is very large

​

  • Using a single minimum support threshold may not be effective

17 of 48

Multiple Minimum Support

  • How to apply multiple minimum supports?
    • MS(i): minimum support for item i
    • e.g.: MS(Milk)=5%, MS(Coke) = 3%,� MS(Broccoli)=0.1%, MS(Salmon)=0.5%
    • MS({Milk, Broccoli}) = min (MS(Milk), MS(Broccoli))� = 0.1%

​

    • Challenge: Support is no longer anti-monotone
      • Suppose: Support(Milk, Coke) = 1.5% and� Support(Milk, Coke, Broccoli) = 0.5%

​

      • {Milk,Coke} is infrequent but {Milk,Coke,Broccoli} is frequent

18 of 48

Multiple Minimum Support

19 of 48

Multiple Minimum Support

20 of 48

Multiple Minimum Support (Liu 1999)

  • Order the items according to their minimum support (in ascending order)
    • e.g.: MS(Milk)=5%, MS(Coke) = 3%,� MS(Broccoli)=0.1%, MS(Salmon)=0.5%
    • Ordering: Broccoli, Salmon, Coke, Milk

​

  • Need to modify Apriori such that:
    • L1 : set of frequent items
    • F1 : set of items whose support is ≥ MS(1)� where MS(1) is mini( MS(i) )
    • C2 : candidate itemsets of size 2 is generated from F1� instead of L1

​

21 of 48

Multiple Minimum Support (Liu 1999)

  • Modifications to Apriori:
    • In traditional Apriori,
      • A candidate (k+1)-itemset is generated by merging two� frequent itemsets of size k
      • The candidate is pruned if it contains any infrequent subsets� of size k
    • Pruning step has to be modified:
      • Prune only if subset contains the first item
      • e.g.: Candidate={Broccoli, Coke, Milk} (ordered according to� minimum support)
      • {Broccoli, Coke} and {Broccoli, Milk} are frequent but � {Coke, Milk} is infrequent
        • Candidate is not pruned because {Coke,Milk} does not contain� the first item, i.e., Broccoli.

22 of 48

Pattern Evaluation

  • Association rule algorithms tend to produce too many rules
    • many of them are uninteresting or redundant
    • Redundant if {A,B,C} → {D} and {A,B} → {D} �have same support & confidence

​

  • Interestingness measures can be used to prune/rank the derived patterns

​

  • In the original formulation of association rules, support & confidence are the only measures used

23 of 48

Application of Interestingness Measure

Interestingness Measures

24 of 48

Computing Interestingness Measure

  • Given a rule X → Y, information needed to compute rule interestingness can be obtained from a contingency table

​

Y

Y

​

X

f11

f10

f1+

X

f01

f00

fo+

​

f+1

f+0

|T|

Contingency table for X → Y

f11: support of X and Y�f10: support of X and Y�f01: support of X and Y�f00: support of X and Y

Used to define various measures

  • support, confidence, lift, Gini,� J-measure, etc.

25 of 48

Drawback of Confidence

​

​

Coffee

​

Coffee

​

Tea

15

5

20

Tea

75

5

80

​

90

10

100

Association Rule: Tea → Coffee�

Confidence= P(Coffee|Tea) = 0.75

but P(Coffee) = 0.9

  • Although confidence is high, rule is misleading
  • P(Coffee|Tea) = 0.9375

26 of 48

Statistical Independence

  • Population of 1000 students
    • 600 students know how to swim (S)
    • 700 students know how to bike (B)
    • 420 students know how to swim and bike (S,B)

​

    • P(S∧B) = 420/1000 = 0.42
    • P(S) × P(B) = 0.6 × 0.7 = 0.42

​

    • P(S∧B) = P(S) × P(B) => Statistical independence
    • P(S∧B) > P(S) × P(B) => Positively correlated
    • P(S∧B) < P(S) × P(B) => Negatively correlated

27 of 48

Statistical-based Measures

  • Measures that take into account statistical dependence

28 of 48

Example: Lift/Interest

​

​

Coffee

​

Coffee

​

Tea

15

5

20

Tea

75

5

80

​

90

10

100

Association Rule: Tea → Coffee�

Confidence= P(Coffee|Tea) = 0.75

but P(Coffee) = 0.9

  • Lift = 0.75/0.9= 0.8333 (< 1, therefore is negatively associated)

29 of 48

Drawback of Lift & Interest

​

Y

Y

​

X

10

0

10

X

0

90

90

​

10

90

100

​

Y

Y

​

X

90

0

90

X

0

10

10

​

90

10

100

Statistical independence:

If P(X,Y)=P(X)P(Y) => Lift = 1

30 of 48

There are lots of measures proposed in the literature

​

Some measures are good for certain applications, but not for others

​

What criteria should we use to determine whether a measure is good or bad?

​

What about Apriori-style support based pruning? How does it affect these measures?

31 of 48

Properties of A Good Measure

  • Piatetsky-Shapiro: �3 properties a good measure M must satisfy:
    • M(A,B) = 0 if A and B are statistically independent

​

    • M(A,B) increase monotonically with P(A,B) when P(A) and P(B) remain unchanged

​

    • M(A,B) decreases monotonically with P(A) [or P(B)] when P(A,B) and P(B) [or P(A)] remain unchanged

32 of 48

Comparing Different Measures

10 examples of contingency tables:

Rankings of contingency tables using various measures:

33 of 48

Property under Variable Permutation

Does M(A,B) = M(B,A)?

​

Symmetric measures:

    • support, lift, collective strength, cosine, Jaccard, etc

Asymmetric measures:

    • confidence, conviction, Laplace, J-measure, etc

34 of 48

Property under Row/Column Scaling

​

Male

Female

​

High

2

3

5

Low

1

4

5

​

3

7

10

​

Male

Female

​

High

4

30

34

Low

2

40

42

​

6

70

76

Grade-Gender Example (Mosteller, 1968):

Mosteller: � Underlying association should be independent of� the relative number of male and female students� in the samples

2x

10x

35 of 48

Property under Inversion Operation

Transaction 1

Transaction N

.

.

.

.

.

36 of 48

Example: φ-Coefficient

  • φ-coefficient is analogous to correlation coefficient for continuous variables

​

Y

Y

​

X

60

10

70

X

10

20

30

​

70

30

100

​

Y

Y

​

X

20

10

30

X

10

60

70

​

30

70

100

φ Coefficient is the same for both tables

37 of 48

Property under Null Addition

Invariant measures:

    • support, cosine, Jaccard, etc

Non-invariant measures:

    • correlation, Gini, mutual information, odds ratio, etc

38 of 48

Different Measures have Different Properties

39 of 48

Support-based Pruning

  • Most of the association rule mining algorithms use support measure to prune rules and itemsets

​

  • Study effect of support pruning on correlation of itemsets
    • Generate 10000 random contingency tables
    • Compute support and pairwise correlation for each table
    • Apply support-based pruning and examine the tables that are removed

40 of 48

Effect of Support-based Pruning

41 of 48

Effect of Support-based Pruning

Support-based pruning eliminates mostly negatively correlated itemsets

42 of 48

Effect of Support-based Pruning

  • Investigate how support-based pruning affects other measures

​

  • Steps:
    • Generate 10000 contingency tables
    • Rank each table according to the different measures
    • Compute the pair-wise correlation between the measures

​

​

43 of 48

Effect of Support-based Pruning

  • Without Support Pruning (All Pairs)
  • Red cells indicate correlation between� the pair of measures > 0.85
  • 40.14% pairs have correlation > 0.85

Scatter Plot between Correlation & Jaccard Measure

44 of 48

Effect of Support-based Pruning

  • 0.5% ≤ support ≤ 50%
  • 61.45% pairs have correlation > 0.85

Scatter Plot between Correlation & Jaccard Measure:

45 of 48

Effect of Support-based Pruning

  • 0.5% ≤ support ≤ 30%
  • 76.42% pairs have correlation > 0.85

Scatter Plot between Correlation & Jaccard Measure

46 of 48

Subjective Interestingness Measure

  • Objective measure:
    • Rank patterns based on statistics computed from data
    • e.g., 21 measures of association (support, confidence, Laplace, Gini, mutual information, Jaccard, etc).

​

  • Subjective measure:
    • Rank patterns according to user’s interpretation
      • A pattern is subjectively interesting if it contradicts the� expectation of a user (Silberschatz & Tuzhilin)
      • A pattern is subjectively interesting if it is actionable� (Silberschatz & Tuzhilin)

47 of 48

Interestingness via Unexpectedness

  • Need to model expectation of users (domain knowledge)

​

​

​

​

​

​

​

​

​

​

  • Need to combine expectation of users with evidence from data (i.e., extracted patterns)

+

Pattern expected to be frequent

-

Pattern expected to be infrequent

Pattern found to be frequent

Pattern found to be infrequent

+

-

Expected Patterns

-

+

Unexpected Patterns

48 of 48

Interestingness via Unexpectedness

  • Web Data (Cooley et al 2001)
    • Domain knowledge in the form of site structure
    • Given an itemset F = {X1, X2, …, Xk} (Xi : Web pages)
      • L: number of links connecting the pages
      • lfactor = L / (k × k-1)
      • cfactor = 1 (if graph is connected), 0 (disconnected graph)
    • Structure evidence = cfactor × lfactor

​

    • Usage evidence

​

    • Use Dempster-Shafer theory to combine domain knowledge and evidence from data