Python实现批量梯度下降不收敛问题求助及代码校验
Hey there! Let's break down why your batch gradient descent isn't converging and fix up your code step by step.
Key Issues in Your Implementation
Let's walk through the problems one by one:
- Learning Rate (α) is way too large: SAT scores are in the 1000-1600 range—using α=0.003 will cause gradient updates to overshoot the cost function's minimum, making the cost explode instead of decrease. You need a tiny learning rate here, like
1e-7or5e-8. - Broken loop condition: Your
while cost > 1line is missing a colon (:) (syntax error), and a fixed threshold like1isn't reliable. Instead, use a maximum iteration limit (to avoid infinite loops) or stop when the cost change between iterations is smaller than a tiny epsilon (e.g.,1e-6). - Incorrect gradient calculation for θ₁: The
matrix_x.transpose() * ...does element-wise multiplication in NumPy, not the dot product needed for the gradient. You should compute the sum ofx * errordirectly, which is the correct gradient term. - Misused global variables: Relying on global variables makes debugging hard and code unmodular. Pass parameters explicitly instead.
- Stale cost value: You only calculate
costonce before the loop—you need to recalculate it every iteration to track convergence.
Fixed Code Implementation
import numpy as np # Assume your dataset is loaded correctly sat = np.array(data['SAT']) gpa = np.array(data['GPA']) m = len(gpa) def calculateCost(theta_0, theta_1, x, y, m): # No global variables—pass all parameters explicitly predictions = theta_0 + theta_1 * x errors = predictions - y cost = (1 / (2 * m)) * np.sum(errors ** 2) return cost def gradDescent(alpha, x, y, m, max_iterations=100000, epsilon=1e-6): theta_0 = 0.01 theta_1 = 0.01 cost_history = [] # Initialize cost current_cost = calculateCost(theta_0, theta_1, x, y, m) cost_history.append(current_cost) for i in range(max_iterations): # Compute predictions and errors predictions = theta_0 + theta_1 * x errors = predictions - y # Calculate gradients correctly grad_0 = (1 / m) * np.sum(errors) grad_1 = (1 / m) * np.sum(x * errors) # Equivalent to dot product # Update parameters temp_0 = theta_0 - alpha * grad_0 temp_1 = theta_1 - alpha * grad_1 theta_0, theta_1 = temp_0, temp_1 # Recalculate cost for this iteration current_cost = calculateCost(theta_0, theta_1, x, y, m) cost_history.append(current_cost) # Check for convergence if abs(cost_history[-1] - cost_history[-2]) < epsilon: print(f"Converged after {i+1} iterations") break return theta_0, theta_1, cost_history # Use a properly sized learning rate alpha = 5e-8 theta_0_final, theta_1_final, cost_history = gradDescent(alpha, sat, gpa, m) print(f"Final theta_0: {theta_0_final:.4f}, Final theta_1: {theta_1_final:.8f}") print(f"Final cost: {cost_history[-1]:.4f}")
Extra Tips for Better Convergence
- Feature Scaling: Standardize your SAT scores (e.g.,
sat_scaled = (sat - np.mean(sat)) / np.std(sat))—this will let you use a larger learning rate and make convergence much faster. - Plot Cost History: After running, plot
cost_historyto verify the cost decreases smoothly over iterations—this is a quick way to confirm your gradient descent is working. - Avoid Globals: By passing parameters explicitly, your code becomes easier to test, reuse, and debug.
内容的提问来源于stack exchange,提问作者dandycheng
相关产品推荐
相关产品推荐

