Learning with Tree-averaged Densities and Distributions
Sergey Kirshner
Alberta Ingenuity Centre for Machine Learning,
Department of Computing Science,
University of Alberta, Canada
December 5, 2007
NIPS 2007
Poster W12
Overview
Learning with Tree-averaged Densities and Distributions
2
NIPS 2007
Most Popular Distribution…
Learning with Tree-averaged Densities and Distributions
3
NIPS 2007
What If the Data Is NOT Gaussian?
Learning with Tree-averaged Densities and Distributions
4
NIPS 2007
Curse of Dimensionality
Learning with Tree-averaged Densities and Distributions
5
NIPS 2007
1/n
1/n
nd cells
V[-2,2]d ≈ 0.9545d
[Bellman 57]
Avoiding the Curse: Step 1�Separating Univariate Marginals
Learning with Tree-averaged Densities and Distributions
6
NIPS 2007
univariate marginals,
independent variables,
multivariate dependence term,
copula
Monotonic Transformation of the Variables
Learning with Tree-averaged Densities and Distributions
7
NIPS 2007
Copula
Learning with Tree-averaged Densities and Distributions
8
NIPS 2007
Copula C is a multivariate distribution (cdf) defined on a unit hypercube with uniform univariate marginals:
Sklar’s Theorem
Learning with Tree-averaged Densities and Distributions
9
NIPS 2007
[Sklar 59]
=
+
Example: Bivariate Gaussian Copula
Learning with Tree-averaged Densities and Distributions
10
NIPS 2007
Useful Properties of Copulas
Learning with Tree-averaged Densities and Distributions
11
NIPS 2007
Copula Density
Learning with Tree-averaged Densities and Distributions
12
NIPS 2007
Separating Univariate Marginals
Learning with Tree-averaged Densities and Distributions
13
NIPS 2007
Inference for the margins [Joe and Xu 96]; canonical maximum likelihood [Genest et al 95]
What Next?
Learning with Tree-averaged Densities and Distributions
14
NIPS 2007
Tree-Structured Densities
Learning with Tree-averaged Densities and Distributions
15
NIPS 2007
x2
x3
x4
x5
x6
x1
Tree-Structured Copulas
Learning with Tree-averaged Densities and Distributions
16
NIPS 2007
Chow-Liu Algorithm (for Copulas)
Learning with Tree-averaged Densities and Distributions
17
NIPS 2007
A1A2
A1A3
A1A4
A2A3
A2A4
A3A4
c(a1,a2)
c(a1,a3)
c(a1,a4)
c(a2,a3)
c(a2,a4)
c(a3,a4)
a1
a3
a2
a4
0.3126
0.0229
0.0172
0.0230
0.0183
0.2603
0.3126
0.0229
0.0172
0.0230
0.0183
0.2603
c(a1,a2)
c(a1,a3)
c(a1,a4)
c(a2,a3)
c(a2,a4)
c(a3,a4)
A1A2
A1A3
A1A4
A2A3
A2A4
A3A4
a1
a3
a2
a4
a1
a3
a2
a4
Distribution over Spanning Trees
Learning with Tree-averaged Densities and Distributions
18
NIPS 2007
a1
a3
a2
a4
β34
β24
β13
β12
β14
β23
a1
a3
a2
a4
β34
β24
β13
β12
β14
β23
[Meilă and Jaakkola 00, 06]
a1
a3
a2
a4
β34
β24
β13
β12
β14
β23
O(d3) !!!
Tree-Averaged Copula
Learning with Tree-averaged Densities and Distributions
19
NIPS 2007
EM for Tree-Averaged Copulas
Learning with Tree-averaged Densities and Distributions
20
NIPS 2007
Intractable!!!
Experiments: Log-Likelihood on Test Data
Learning with Tree-averaged Densities and Distributions
21
NIPS 2007
UCI ML Repository
MAGIC data set
12000 10-dimensional vectors
2000 examples in test sets
Average over 10 partitions
Binary-Continuous Data
Learning with Tree-averaged Densities and Distributions
22
NIPS 2007
Summary
Learning with Tree-averaged Densities and Distributions
23
NIPS 2007