如何定义分布上的核矩阵?结合HSIC相关讲义内容解析
Got it, let's walk through this clearly—first, we'll define what a kernel matrix over a distribution is, then tie that directly to the HSIC details you shared.
1. Defining Kernel Matrices Over Distributions
Let's start with the sample-level kernel matrix you're seeing in the HSIC formula, since that's the easiest way to bridge to the distribution version.
- For a set of i.i.d samples $x_1, x_2, ..., x_n$ from random variable $X$'s distribution, the sample kernel matrix $K$ is an $n \times n$ symmetric matrix where every entry is:
$$K_{ij} = k(x_i, x_j)$$
Here, $k(\cdot, \cdot)$ is a valid kernel function (think RBF, polynomial, linear) that meets Mercer's condition—this lets us implicitly map data into a high-dimensional Hilbert space without calculating the actual mapping, which is the magic of kernel methods. - Now, the distribution-level kernel matrix (or more formally, the kernel operator associated with the distribution) is what you get when you move from finite samples to the entire distribution $\mathcal{P}_X$. Instead of computing similarities between individual samples, we use expectations over the distribution. For two independent copies $X, X' \sim \mathcal{P}_X$, the core distribution-level kernel term is $\mathbb{E}[k(X, X')]$. You can think of this as the "average similarity" between any two points drawn from $X$'s distribution. In practice, when we talk about HSIC at the theoretical (distribution) level, we use these expectation-based kernel terms to define the population HSIC.
2. Linking to HSIC (From Your Lecture Notes)
Let's connect this directly to the HSIC content you provided:
希尔伯特-施密特独立性准则(HSIC)用于衡量两个随机变量$X$和$Y$的依赖关系。HSIC的经验估计值与$trace(KHLH)$成正比,其中$K$是$X$的核矩阵,$L$是$Y$的核矩阵,$H$是中心化矩阵,满足$H_{ij} = δ(i, j) − rac{1}{n}$。当且仅当$X$和$Y$相互独立时,$HSIC(X, Y ) = 0$。$HSIC(X, Y)$的值越大,$X$与$Y$之间的依赖关系越强……
Breaking Down the Empirical HSIC
The empirical HSIC (the $trace(KHLH)$ term scaled appropriately) is an estimate of the population (distribution-level) HSIC. Here's what each part does:
- $K$ and $L$ are the sample kernel matrices for $X$ and $Y$ respectively—they capture pairwise similarities between your observed samples.
- $H$ is the centering matrix: $H_{ij} = \delta(i,j) - 1/n$. When you multiply $K$ by $H$ (i.e., $HKH$), you're centering the kernel matrix—this removes the average similarity across all samples, which is equivalent to centering the mapped data in the Hilbert space.
- The trace of $KHLH$ (written as $\text{trace}(KHLH)$) computes the sum of the diagonal entries of the resulting matrix. This scalar value measures the covariance between the centered kernel representations of $X$ and $Y$—the more aligned these representations are, the stronger the dependence between $X$ and $Y$.
Key Properties (From Your Notes)
- Independence Check: $\text{HSIC}(X,Y) = 0$ if and only if $X$ and $Y$ are independent. This holds exactly at the distribution level, and asymptotically (as sample size grows) for the empirical estimate.
- Dependency Strength: Larger HSIC values mean stronger dependence between $X$ and $Y$. This makes intuitive sense—if the centered similarity patterns of $X$ samples match those of $Y$ samples closely, the variables must be linked.
3. A Quick Concrete Example
Let's say we use an RBF kernel for $X$: $k(x_i, x_j) = \exp(-||x_i - x_j||2/(2\sigma2))$, which measures how "close" each pair of $X$ samples is. For $Y$, we use a linear kernel: $l(y_i, y_j) = y_i^T y_j$.
- The sample kernel matrix $K$ will have high values for pairs of $X$ samples that are close together, low values otherwise.
- $HKH$ adjusts $K$ so that the average value across all entries is zero (centering it).
- $\text{trace}(KHLH)$ then tells us how much this centered similarity structure of $X$ overlaps with the centered similarity structure of $Y$. If $X$ and $Y$ are independent, this overlap is zero; if $Y$ increases whenever $X$ does, this value will be large.
内容的提问来源于stack exchange,提问作者mercury0114

