基于牛顿法的梯度下降模型:损失函数Hessian矩阵求解问询
Got it, let's break down how to compute the Hessian matrix for your given loss function step by step. I'll start by restating key definitions and work through the second partial derivatives clearly.
Quick Recap of Key Equations
First, let's refresh our memory with your loss function and first-order partial derivative, using simplified notation to make calculations smoother:
Loss function:
$$f(w) = \frac{1}{2}|w|^2 + C \sum_{i=1}^n\log(1 + \exp(-y_iw^Tx_i))$$
First partial derivative for $w_j$ (using the sigmoid function $\sigma(z) = \frac{\exp(z)}{1+\exp(z)}$, where $\sigma(-y_iw^Tx_i) = \frac{\exp(-y_iwTx_i)}{1+\exp(-y_iwTx_i)}$):
$$\frac{\partial f(w)}{\partial w_j} = w_j - C \sum_{i=1}^n x_{ij}y_i \sigma(-y_iw^Tx_i)$$
A critical sigmoid property we’ll rely on:
$$\sigma'(z) = \sigma(z)(1 - \sigma(z))$$
Calculating the Hessian Elements
The Hessian matrix $H$ has elements $H_{jk} = \frac{\partial^2 f}{\partial w_j \partial w_k}$. Let's compute this for two cases: $j = k$ and $j \neq k$.
Case 1: $j \neq k$
Differentiate the first partial derivative with respect to $w_k$:
$$\frac{\partial^2 f}{\partial w_j \partial w_k} = \frac{\partial}{\partial w_k}\left(w_j\right) - C \sum_{i=1}^n x_{ij} y_i \frac{\partial}{\partial w_k}\left(\sigma(-y_i w^T x_i)\right)$$
- The first term $\frac{\partial}{\partial w_k}(w_j) = 0$ (since $j$ and $k$ are distinct weights).
- For the second term, let $z_i = -y_i w^T x_i$. Using the chain rule:
$$\frac{\partial \sigma(z_i)}{\partial w_k} = \sigma'(z_i) \cdot \frac{\partial z_i}{\partial w_k} = \sigma(z_i)(1 - \sigma(z_i)) \cdot (-y_i x_{ik})$$
Substitute back into the sum and simplify: the negative signs cancel, and $y_i^2 = 1$ (since $y_i$ is a classification label, typically $\pm1$). We get:
$$H_{jk} = C \sum_{i=1}^n x_{ij} x_{ik} \sigma(-y_i w^T x_i)(1 - \sigma(-y_i w^T x_i))$$
Case 2: $j = k$
When $j = k$, the first term $\frac{\partial}{\partial w_j}(w_j) = 1$. The rest of the calculation matches the $j \neq k$ case, so we add this 1 to the sum:
$$H_{jj} = 1 + C \sum_{i=1}^n x_{ij}^2 \sigma(-y_i w^T x_i)(1 - \sigma(-y_i w^T x_i))$$
Compact Matrix Form of the Hessian
For practical use in Newton's method, we can write the Hessian in matrix notation:
- Let $X$ be the $n \times d$ feature matrix (each row is $x_i^T$)
- Let $D$ be an $n \times n$ diagonal matrix where $D_{ii} = \sigma(-y_i w^T x_i)(1 - \sigma(-y_i w^T x_i))$
The Hessian becomes:
$$H = I_d + C X^T D X$$
where $I_d$ is the $d \times d$ identity matrix.
This form is great because $H$ is positive definite (we add the identity to a positive semi-definite matrix), which ensures Newton's method will converge reliably.
内容的提问来源于stack exchange,提问作者kjl

