修改单隐藏层神经网络为单浮点输出后反向传播梯度异常求助
Hey there! Let's break down where your backprop code is going wrong—these are easy mistakes to make when adapting a classification network to regression with a linear output layer.
Key Issues in Your Current Code
Let's walk through the critical errors step by step:
Incorrect delta3 calculation
You setdelta3 = (a3-y(i))^2;, which is the squared error for the sample, not the gradient of the cost function with respect toz3. For a linear output layer (a3 = z3) and mean squared error cost:- The cost for a single sample is
(a3 - y(i))²(before averaging over all samples) - The gradient of this cost with respect to
z3is2*(a3 - y(i))(sinceda3/dz3 = 1for linear activation)
Using the squared error instead of the first derivative completely breaks the gradient signal here.
- The cost for a single sample is
Missing activation function derivative in delta2
Your hidden layer usestanhactivation, so when computingdelta2(the gradient with respect toz2), you need to multiply by the derivative oftanh(z2)—which is1 - tanh(z2)^2(or equivalently1 - a2(2:end).^2, sincea2 = tanh(z2)). Skipping this step throws off all hidden layer gradient calculations.Incorrect delta2 weight multiplication
When computingdelta2, you should exclude the bias term fromTheta2(the bias node ina2doesn't correspond to anyz2value). You need to useTheta2(:, 2:end)instead of the fullTheta2'when multiplying bydelta3.
Corrected Backpropagation Code
Here's the fixed version of your backprop loop, with comments explaining each change:
% Initialize gradients (don't forget this step before the loop!) Theta1_grad = zeros(size(Theta1)); Theta2_grad = zeros(size(Theta2)); %反向传播,遍历m个样本累积梯度后取平均: for i = 1:m a1 = [1; X(i, :)']; %修正为列向量以匹配矩阵乘法维度 z2 = Theta1 * a1; a2 = tanh(z2); a2 = [1; a2]; % 添加偏置节点 z3 = Theta2 * a2; a3 = z3; % 线性输出 % 修正delta3:代价对z3的偏导 delta3 = 2 * (a3 - y(i)); % 修正delta2:去掉Theta2中对应a2偏置的列,再乘tanh的导数 tanh_derivative = 1 - a2(2:end).^2; delta2 = (Theta2(:, 2:end)' * delta3) .* tanh_derivative; % 累积梯度(确保矩阵维度匹配) Theta2_grad = Theta2_grad + delta3 * a2'; Theta1_grad = Theta1_grad + delta2 * a1'; endfor Theta1_grad = (1/m) * Theta1_grad; % 对所有样本取平均 Theta2_grad = (1/m) * Theta2_grad;
Quick Tips to Verify Gradients
To confirm your gradients are correct, compute numerical gradients for comparison:
- For each parameter in
Theta1andTheta2, add a tiny epsilon (like1e-4), calculate the cost, subtract the cost when the parameter is reduced by epsilon, then divide by2*epsilonto get a numerical gradient approximation. - Your corrected analytical gradients should match the numerical ones very closely (within
1e-3or better).
Optional Cost Function Tweak
Your current cost function is correct, but adding a factor of 0.5 (Cost = 0.5*sum((y-a3).^2)/m;) cancels out the factor of 2 in the gradient, making the math cleaner without changing the optimization direction.
内容的提问来源于stack exchange,提问作者aletherios

