You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现批量梯度下降不收敛问题求助及代码校验

Hey there! Let's break down why your batch gradient descent isn't converging and fix up your code step by step.

Key Issues in Your Implementation

Let's walk through the problems one by one:

  • Learning Rate (α) is way too large: SAT scores are in the 1000-1600 range—using α=0.003 will cause gradient updates to overshoot the cost function's minimum, making the cost explode instead of decrease. You need a tiny learning rate here, like 1e-7 or 5e-8.
  • Broken loop condition: Your while cost > 1 line is missing a colon (:) (syntax error), and a fixed threshold like 1 isn't reliable. Instead, use a maximum iteration limit (to avoid infinite loops) or stop when the cost change between iterations is smaller than a tiny epsilon (e.g., 1e-6).
  • Incorrect gradient calculation for θ₁: The matrix_x.transpose() * ... does element-wise multiplication in NumPy, not the dot product needed for the gradient. You should compute the sum of x * error directly, which is the correct gradient term.
  • Misused global variables: Relying on global variables makes debugging hard and code unmodular. Pass parameters explicitly instead.
  • Stale cost value: You only calculate cost once before the loop—you need to recalculate it every iteration to track convergence.

Fixed Code Implementation

import numpy as np

# Assume your dataset is loaded correctly
sat = np.array(data['SAT'])
gpa = np.array(data['GPA'])
m = len(gpa)

def calculateCost(theta_0, theta_1, x, y, m):
    # No global variables—pass all parameters explicitly
    predictions = theta_0 + theta_1 * x
    errors = predictions - y
    cost = (1 / (2 * m)) * np.sum(errors ** 2)
    return cost

def gradDescent(alpha, x, y, m, max_iterations=100000, epsilon=1e-6):
    theta_0 = 0.01
    theta_1 = 0.01
    cost_history = []
    
    # Initialize cost
    current_cost = calculateCost(theta_0, theta_1, x, y, m)
    cost_history.append(current_cost)
    
    for i in range(max_iterations):
        # Compute predictions and errors
        predictions = theta_0 + theta_1 * x
        errors = predictions - y
        
        # Calculate gradients correctly
        grad_0 = (1 / m) * np.sum(errors)
        grad_1 = (1 / m) * np.sum(x * errors)  # Equivalent to dot product
        
        # Update parameters
        temp_0 = theta_0 - alpha * grad_0
        temp_1 = theta_1 - alpha * grad_1
        
        theta_0, theta_1 = temp_0, temp_1
        
        # Recalculate cost for this iteration
        current_cost = calculateCost(theta_0, theta_1, x, y, m)
        cost_history.append(current_cost)
        
        # Check for convergence
        if abs(cost_history[-1] - cost_history[-2]) < epsilon:
            print(f"Converged after {i+1} iterations")
            break
    
    return theta_0, theta_1, cost_history

# Use a properly sized learning rate
alpha = 5e-8
theta_0_final, theta_1_final, cost_history = gradDescent(alpha, sat, gpa, m)

print(f"Final theta_0: {theta_0_final:.4f}, Final theta_1: {theta_1_final:.8f}")
print(f"Final cost: {cost_history[-1]:.4f}")

Extra Tips for Better Convergence

  • Feature Scaling: Standardize your SAT scores (e.g., sat_scaled = (sat - np.mean(sat)) / np.std(sat))—this will let you use a larger learning rate and make convergence much faster.
  • Plot Cost History: After running, plot cost_history to verify the cost decreases smoothly over iterations—this is a quick way to confirm your gradient descent is working.
  • Avoid Globals: By passing parameters explicitly, your code becomes easier to test, reuse, and debug.

内容的提问来源于stack exchange,提问作者dandycheng

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:15:06