反向传播优化:如何利用导数调整权重与偏置以最小化代价函数?
Great question—let’s break this down step by step, since this is the core of how neural networks learn from data.
1. 导数(梯度)在参数优化中的核心作用
First, let’s get the intuition straight: your cost function (like MSE for regression or cross-entropy for classification) tells you how far your model’s predictions are from the true labels. To minimize this cost, you need to adjust the weights and biases of your network.
The partial derivative of the cost function with respect to a weight (or bias) tells you two key things:
- Direction: Whether increasing that parameter will raise or lower the cost. If
∂C/∂wis positive, boostingwmakes the cost go up; if negative, boostingwbrings the cost down. - Magnitude: How sensitive the cost is to changes in that parameter. A large absolute value means the parameter has a big impact on the cost right now.
Backpropagation is just the efficient way to compute all these partial derivatives for every weight and bias in the network. Once you have those gradients, you know exactly which way to tweak each parameter to shrink the cost.
2. 已知梯度时的参数调整逻辑
Yes, the core idea is to adjust each parameter by moving it in the opposite direction of the gradient—because the gradient points to where the cost increases the fastest. The basic update formula looks like this:
# 更新权重 w = w - learning_rate * dC_dw # 更新偏置 b = b - learning_rate * dC_db
Here’s why this works: if dC_dw is positive, subtracting learning_rate * dC_dw decreases w, which pushes the cost down. If dC_dw is negative, subtracting that term increases w—again, pushing the cost down. The learning rate (learning_rate) is that constant you’re asking about—it controls how big a step you take in the gradient’s opposite direction.
3. 只是“减导数乘常数”吗?
The pure batch gradient descent algorithm does exactly this, using the average gradient over your entire training dataset. But in practice, we rarely use this vanilla version because it’s slow for large datasets and can get stuck in local minima.
Instead, we use variations that build on this core logic:
- Stochastic Gradient Descent (SGD): Updates parameters using the gradient from a single training sample instead of the whole batch. This adds randomness that can help escape local minima, though it makes the loss curve noisier.
- Mini-batch Gradient Descent: Uses a small subset (mini-batch) of the data to compute the average gradient—balancing the stability of batch GD and the speed/randomness of SGD.
- Momentum: Adds a "velocity" term that accumulates past gradients, reducing oscillations and helping the model converge faster, especially in steep or narrow valleys of the cost function.
- Adam/RMSprop: These adaptive optimizers adjust the learning rate for each parameter individually based on past gradients, so you don’t have to tune a single global learning rate as carefully.
All of these methods still rely on moving parameters against the gradient—they just tweak how the gradient is used to make training more efficient and stable.
4. 如何选择学习率(那个“常数”)
The learning rate is one of the most critical hyperparameters to get right—here’s how to approach it:
- Start small: Common initial values are
1e-4,1e-3, or1e-2(stick to the smaller end if you’re unsure, especially for large models). - Watch the loss curve:
- If your loss is bouncing wildly (large oscillations), your learning rate is too big—you’re overshooting the minimum.
- If your loss is barely decreasing, your learning rate is too small—you’re taking tiny steps and won’t converge in reasonable time.
- Use learning rate scheduling: Gradually reduce the learning rate as training progresses (e.g., multiply by 0.1 every 10 epochs, or use cosine annealing). This lets you take big steps early to find a good parameter space, then smaller steps to fine-tune around the minimum.
- Try adaptive optimizers: Adam, RMSprop, or AdaGrad automatically adjust the learning rate for each parameter, so you don’t have to manually tune it as much. They’re a great default for most tasks.
内容的提问来源于stack exchange,提问作者Ron Lauterbach

