You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于Witten《数据挖掘》中反向传播向量矩阵形式方程推导的疑问

Clarifying the Cross-Entropy Loss Derivative with Softmax (Ian Witten's Data Mining 2016)

Hey there, let's walk through this derivation step by step to clear up the confusion between your result and the one in the book. First, let's make sure we're starting with the same loss function (I'll clean up the formatting a bit for clarity):

The original loss function from page 427 is:
$$L = -\sum_{k=1}^K y_k \left[ a_k - \log\left( \sum_{c=1}^K \exp(a_c) \right) \right]$$

Key Background: One-Hot Label Assumption

First, a critical detail the book relies on (standard for classification tasks) is that $y_k$ is a one-hot encoded vector: for the true class $j$, $y_j = 1$, and all other $y_k = 0$ (where $k \neq j$). This is why the book's derivation simplifies the sum—let's see how.

Step-by-Step Derivation

Let's split the loss function into two separate terms to make differentiation easier:
$$L = -\sum_{k=1}^K y_k a_k + \sum_{k=1}^K y_k \log\left( \sum_{c=1}^K \exp(a_c) \right)$$

Now compute $\frac{\partial L}{\partial a_j}$ by differentiating each term:

  1. First Term Derivative:
    $$\frac{\partial}{\partial a_j} \left( -\sum_{k=1}^K y_k a_k \right) = -y_j$$
    Only the $k=j$ term contributes here (since $\frac{\partial a_k}{\partial a_j} = 1$ if $k=j$, 0 otherwise), and all other $y_k$ are 0 anyway due to one-hot encoding.

  2. Second Term Derivative:
    The log term $\log\left( \sum_{c=1}^K \exp(a_c) \right)$ doesn't depend on $k$, so we can rewrite the sum as:
    $$\left( \sum_{k=1}^K y_k \right) \cdot \log\left( \sum_{c=1}^K \exp(a_c) \right)$$
    Since $\sum_{k=1}^K y_k = 1$ (one-hot labels sum to 1), this simplifies to just $\log\left( \sum_{c=1}^K \exp(a_c) \right)$. Now differentiate with respect to $a_j$:
    $$\frac{\partial}{\partial a_j} \log\left( \sum_{c=1}^K \exp(a_c) \right) = \frac{\exp(a_j)}{\sum_{c=1}^K \exp(a_c)}$$

Combine the Two Terms

Adding the derivatives together gives:
$$\frac{\partial L}{\partial a_j} = -y_j + \frac{\exp(a_j)}{\sum_{c=1}^K \exp(a_c)} = -\left( y_j - \frac{\exp(a_j)}{\sum_{c=1}^K \exp(a_c)} \right)$$
Which matches exactly what's in the book.

Why Your Derivation Looks Different (But Is Equivalent)

Your result is:
$$- \sum_{k=1}^K y_k \left[ \mathbb{1}{k=j} - \frac{\exp(a_j)}{\sum{c=1}^K \exp(a_c)} \right]$$
Let's expand this sum using the one-hot property:

  • For $k=j$: $y_j = 1$ and $\mathbb{1}_{k=j} = 1$, so this term is $-1 \cdot \left(1 - \frac{\exp(a_j)}{\sum \exp(a_c)}\right)$
  • For $k \neq j$: $y_k = 0$, so all these terms vanish

Expanding the non-zero term gives:
$$-1 + \frac{\exp(a_j)}{\sum \exp(a_c)} = -\left(1 - \frac{\exp(a_j)}{\sum \exp(a_c)}\right) = -\left(y_j - \frac{\exp(a_j)}{\sum \exp(a_c)}\right)$$
Which is identical to the book's result!

Summary

  • Your derivation is mathematically equivalent to the book's—you just kept the sum intact, while the book used the one-hot label property to simplify it early on.
  • The key insight is leveraging the one-hot nature of $y_k$ to eliminate all non-relevant terms in the sum, which is a standard shortcut in neural network loss derivations.

内容的提问来源于stack exchange,提问作者Roulbacha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 06:23:51