Data Mining �Association Analysis: FP-Growth algorithm
FP-growth Algorithm
FP-tree construction
null
A:1
B:1
null
A:1
B:1
B:1
C:1
D:1
After reading TID=1:
After reading TID=2:
FP-Tree Construction
null
A:7
B:5
B:3
C:3
D:1
C:1
D:1
C:3
D:1
D:1
E:1
E:1
Pointers are used to assist frequent itemset generation
D:1
E:1
Transaction Database
Header table
FP-growth
null
A:7
B:5
B:1
C:1
D:1
C:1
D:1
C:3
D:1
D:1
Conditional Pattern base for D: � P = {(A:1,B:1,C:1),� (A:1,B:1), � (A:1,C:1),� (A:1), � (B:1,C:1)}
Recursively apply FP-growth on P
Frequent Itemsets found (with sup > 1):� AD, BD, CD, ACD, BCD
D:1
Tree Projection
Set enumeration tree:
Possible Extension: E(A) = {B,C,D,E}
Possible Extension: E(ABC) = {D,E}
Tree Projection
Projected Database
Original Database:
Projected Database for node A:
For each transaction T, projected transaction at node A is T ∩ E(A)
ECLAT
TID-list
ECLAT
∧
→
Rule Generation
ABC →D, ABD →C, ACD →B, BCD →A, �A →BCD, B →ACD, C →ABD, D →ABC�AB →CD, AC → BD, AD → BC, BC →AD, �BD →AC, CD →AB, �
Rule Generation
c(ABC →D) can be larger or smaller than c(AB →D)
Rule Generation for Apriori Algorithm
Lattice of rules
Pruned Rules
Low Confidence Rule
Rule Generation for Apriori Algorithm
Effect of Support Distribution
Support distribution of a retail data set
Effect of Support Distribution
Multiple Minimum Support
Multiple Minimum Support
Multiple Minimum Support
Multiple Minimum Support (Liu 1999)
Multiple Minimum Support (Liu 1999)
Pattern Evaluation
Application of Interestingness Measure
Interestingness Measures
Computing Interestingness Measure
| Y | Y | |
X | f11 | f10 | f1+ |
X | f01 | f00 | fo+ |
| f+1 | f+0 | |T| |
Contingency table for X → Y
f11: support of X and Y�f10: support of X and Y�f01: support of X and Y�f00: support of X and Y
Used to define various measures
Drawback of Confidence
| Coffee | Coffee | |
Tea | 15 | 5 | 20 |
Tea | 75 | 5 | 80 |
| 90 | 10 | 100 |
Association Rule: Tea → Coffee�
Confidence= P(Coffee|Tea) = 0.75
but P(Coffee) = 0.9
Statistical Independence
Statistical-based Measures
Example: Lift/Interest
| Coffee | Coffee | |
Tea | 15 | 5 | 20 |
Tea | 75 | 5 | 80 |
| 90 | 10 | 100 |
Association Rule: Tea → Coffee�
Confidence= P(Coffee|Tea) = 0.75
but P(Coffee) = 0.9
Drawback of Lift & Interest
| Y | Y | |
X | 10 | 0 | 10 |
X | 0 | 90 | 90 |
| 10 | 90 | 100 |
| Y | Y | |
X | 90 | 0 | 90 |
X | 0 | 10 | 10 |
| 90 | 10 | 100 |
Statistical independence:
If P(X,Y)=P(X)P(Y) => Lift = 1
There are lots of measures proposed in the literature
Some measures are good for certain applications, but not for others
What criteria should we use to determine whether a measure is good or bad?
What about Apriori-style support based pruning? How does it affect these measures?
Properties of A Good Measure
Comparing Different Measures
10 examples of contingency tables:
Rankings of contingency tables using various measures:
Property under Variable Permutation
Does M(A,B) = M(B,A)?
Symmetric measures:
Asymmetric measures:
Property under Row/Column Scaling
| Male | Female | |
High | 2 | 3 | 5 |
Low | 1 | 4 | 5 |
| 3 | 7 | 10 |
| Male | Female | |
High | 4 | 30 | 34 |
Low | 2 | 40 | 42 |
| 6 | 70 | 76 |
Grade-Gender Example (Mosteller, 1968):
Mosteller: � Underlying association should be independent of� the relative number of male and female students� in the samples
2x
10x
Property under Inversion Operation
Transaction 1
Transaction N
.
.
.
.
.
Example: φ-Coefficient
| Y | Y | |
X | 60 | 10 | 70 |
X | 10 | 20 | 30 |
| 70 | 30 | 100 |
| Y | Y | |
X | 20 | 10 | 30 |
X | 10 | 60 | 70 |
| 30 | 70 | 100 |
φ Coefficient is the same for both tables
Property under Null Addition
Invariant measures:
Non-invariant measures:
Different Measures have Different Properties
Support-based Pruning
Effect of Support-based Pruning
Effect of Support-based Pruning
Support-based pruning eliminates mostly negatively correlated itemsets
Effect of Support-based Pruning
Effect of Support-based Pruning
Scatter Plot between Correlation & Jaccard Measure
Effect of Support-based Pruning
Scatter Plot between Correlation & Jaccard Measure:
Effect of Support-based Pruning
Scatter Plot between Correlation & Jaccard Measure
Subjective Interestingness Measure
Interestingness via Unexpectedness
+
Pattern expected to be frequent
-
Pattern expected to be infrequent
Pattern found to be frequent
Pattern found to be infrequent
+
-
Expected Patterns
-
+
Unexpected Patterns
Interestingness via Unexpectedness