反向传播中神经元误差对网络参数(权重、偏置)的影响问询
Great questions—these are core mechanics that often get glossed over in introductory courses, so let’s unpack them step by step.
1. How do neuron errors translate to weight/bias adjustments?
First, let’s clarify what that "neuron error" actually represents: it’s the gradient of the loss function with respect to the neuron’s weighted input (often denoted as δ_j = ∂L/∂z_j, where z_j is the sum of inputs times weights plus bias for neuron j). This error tells us how much changing z_j would affect the overall loss—our goal is to tweak weights and biases to reduce this loss.
Here’s the direct link between error and updates, using gradient descent:
- For a weight
w_ijthat connects neuroni(from the previous layer) to neuronj(current layer):
The gradient of the loss with respect tow_ijis∂L/∂w_ij = δ_j * a_i, wherea_iis the output activation of neuroni.
We update the weight using this gradient:w_ij = w_ij - α * (δ_j * a_i)(α is the learning rate). - For a bias
b_jof neuronj:
The gradient is∂L/∂b_j = δ_j(since the bias’s "input" is always 1).
Update the bias:b_j = b_j - α * δ_j.
In short: the neuron error δ_j is the key term that scales how much we adjust every weight feeding into j (paired with the input activation from the previous layer) and the bias of j itself.
2. How does a single training sample’s neuron error affect the entire network?
When you run backpropagation on a single sample (this is called Stochastic Gradient Descent, SGD), the error you compute for each neuron is specific to that sample’s loss. This error propagates backward through every layer, and every weight in the network gets an update based on that single sample’s gradient.
Let’s walk through a tiny example to make this concrete:
Suppose we have a 3-layer network: Input → Hidden → Output.
- Forward pass: Compute activations for hidden and output layers using the current weights.
- Calculate output layer error
δ_outputusing the sample’s true label and the derivative of the output activation function. - Use
δ_outputto compute the hidden layer errorδ_hidden:δ_hidden = δ_output * w_hidden_to_output * σ’(z_hidden)(σ’ is the derivative of the hidden layer’s activation function). - Update every weight in the network:
- Weights from Hidden to Output: Updated using
δ_outputand hidden layer activations. - Weights from Input to Hidden: Updated using
δ_hiddenand input layer values.
- Weights from Hidden to Output: Updated using
- Update both hidden and output layer biases using their respective errors.
The key point: a single sample’s error doesn’t just affect one part of the network—it ripples backward, driving updates for every weight and bias that contributes to the sample’s prediction. This is why SGD updates are "noisy" (each update is based on one sample’s signal) but can be fast to converge, especially with large datasets.
内容的提问来源于stack exchange,提问作者Prze Gee

