反向传播原理及无隐藏层简单神经网络的权重调整方法咨询
Let's break this down step by step—since you've got a no-hidden-layer network, this is actually just a linear regression model, and backpropagation here ties directly to gradient descent in a super straightforward way.
First, Define Your Model
Your network takes 5 inputs (x = [x_1, x_2, x_3, x_4, x_5]) and outputs a weighted sum (since there's no activation function or hidden layer):
(y_{pred} = w_1x_1 + w_2x_2 + w_3x_3 + w_4x_4 + w_5x_5)
(If you had a bias term, it would be (y_{pred} = \sum w_i x_i + b), but we'll stick with your setup for now.)
We'll use the Mean Squared Error (MSE) loss function, the most common choice for regression tasks:
(L = \frac{1}{2}(y_{true} - y_{pred})^2)
(The 1/2 is just a trick to make differentiation cleaner—it doesn't change the direction of our updates.)
Calculating Partial Derivatives (The "Why" Behind Weight Adjustments)
You mentioned confusion about choosing which partial derivatives to compute. For backpropagation, we need the derivative of the loss with respect to each weight—this tells us how much changing that weight will affect the total loss.
Using the chain rule, here's the breakdown for a single weight (w_i):
[
\frac{\partial L}{\partial w_i} = -(y_{true} - y_{pred}) \cdot x_i
]
In plain English: the partial derivative equals the negative of your prediction error multiplied by the corresponding input value. That's it! This is the gradient we need for gradient descent.
Adjusting Weights with Gradient Descent
The weight update formula is:
(w_i^{new} = w_i^{old} - \alpha \cdot \frac{\partial L}{\partial w_i})
Where (\alpha) is the learning rate (a small positive number like 0.01 or 0.001 that controls how big each weight step is).
Let's Plug in Your Example
You have:
- (y_{pred} = 8), (y_{true} = 8.5) → error = (8.5 - 8 = 0.5)
- Initial weights: ([0.2, 0.2, 0.15, 0.15, 0.3])
Suppose your input was, say, (x = [2, 3, 1, 4, 5]) (the exact values don't change the logic). For each weight:
- For (w_1): (\frac{\partial L}{\partial w_1} = -0.5 * 2 = -1)
Update: (0.2 - \alpha*(-1) = 0.2 + \alpha*1) - For (w_2): (\frac{\partial L}{\partial w_2} = -0.5 3 = -1.5)
Update: (0.2 + \alpha1.5)
If (\alpha = 0.01), (w_1) becomes 0.21, (w_2) becomes 0.215, and so on. Each adjustment pushes the next prediction closer to 8.5.
Why This Partial Derivative?
We're trying to minimize the loss function. Gradient descent works by moving weights in the direction that reduces loss the fastest—that direction is the negative of the gradient (hence the minus sign in the update formula). By taking the derivative of loss with respect to each weight, we isolate exactly how much each weight contributes to the error, so we can tweak each one individually to fix the prediction.
Easy-to-Understand English Resources
- Neural Networks and Deep Learning by Michael Nielsen: A free, beginner-friendly book that walks through backpropagation with simple math and code examples—starts with basic perceptrons and builds up slowly.
- Machine Learning Yearning by Andrew Ng: While not focused solely on backpropagation, it uses plain language to explain core concepts like gradient descent and error minimization, which will help solidify your intuition.
- The "Linear Layer Backpropagation" section in Deep Learning by Goodfellow, Bengio, and Courville: This is a textbook staple, but the linear layer part is straightforward and covers the exact scenario you're working with.
内容的提问来源于stack exchange,提问作者Yehor Naumenko

