You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

反向传播中神经元误差对网络参数(权重、偏置)的影响问询

Understanding How Backpropagation Errors Drive Weight Updates

Great questions—these are core mechanics that often get glossed over in introductory courses, so let’s unpack them step by step.


1. How do neuron errors translate to weight/bias adjustments?

First, let’s clarify what that "neuron error" actually represents: it’s the gradient of the loss function with respect to the neuron’s weighted input (often denoted as δ_j = ∂L/∂z_j, where z_j is the sum of inputs times weights plus bias for neuron j). This error tells us how much changing z_j would affect the overall loss—our goal is to tweak weights and biases to reduce this loss.

Here’s the direct link between error and updates, using gradient descent:

  • For a weight w_ij that connects neuron i (from the previous layer) to neuron j (current layer):
    The gradient of the loss with respect to w_ij is ∂L/∂w_ij = δ_j * a_i, where a_i is the output activation of neuron i.
    We update the weight using this gradient: w_ij = w_ij - α * (δ_j * a_i) (α is the learning rate).
  • For a bias b_j of neuron j:
    The gradient is ∂L/∂b_j = δ_j (since the bias’s "input" is always 1).
    Update the bias: b_j = b_j - α * δ_j.

In short: the neuron error δ_j is the key term that scales how much we adjust every weight feeding into j (paired with the input activation from the previous layer) and the bias of j itself.


2. How does a single training sample’s neuron error affect the entire network?

When you run backpropagation on a single sample (this is called Stochastic Gradient Descent, SGD), the error you compute for each neuron is specific to that sample’s loss. This error propagates backward through every layer, and every weight in the network gets an update based on that single sample’s gradient.

Let’s walk through a tiny example to make this concrete:
Suppose we have a 3-layer network: Input → Hidden → Output.

  1. Forward pass: Compute activations for hidden and output layers using the current weights.
  2. Calculate output layer error δ_output using the sample’s true label and the derivative of the output activation function.
  3. Use δ_output to compute the hidden layer error δ_hidden: δ_hidden = δ_output * w_hidden_to_output * σ’(z_hidden) (σ’ is the derivative of the hidden layer’s activation function).
  4. Update every weight in the network:
    • Weights from Hidden to Output: Updated using δ_output and hidden layer activations.
    • Weights from Input to Hidden: Updated using δ_hidden and input layer values.
  5. Update both hidden and output layer biases using their respective errors.

The key point: a single sample’s error doesn’t just affect one part of the network—it ripples backward, driving updates for every weight and bias that contributes to the sample’s prediction. This is why SGD updates are "noisy" (each update is based on one sample’s signal) but can be fast to converge, especially with large datasets.


内容的提问来源于stack exchange,提问作者Prze Gee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:58:12