You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Gradient Descent拟合异常:散点负相关但拟合线正相关求助

梯度下降拟合线与数据趋势相反的问题排查与解决

问题根源分析

你的拟合线呈现正相关但数据实际为负相关,核心原因是特征未做归一化+学习率过小+迭代次数不足,导致梯度下降在有限迭代内无法收敛到正确参数:

  • Boston数据集的lstat(低收入人群占比)范围约0-37,medv(房价)范围约5-50,两者尺度差异极大,未归一化时梯度量级会异常大,小学习率下参数更新极慢。
  • 初始参数w=0时,梯度计算结果为负(prediction - y为负,乘以正的x后求和平均为负),按梯度下降规则w -= alpha * w_gradient,会让w往正方向小幅增加,但真实最优w是负数,小学习率+少迭代次数下,参数还未走到正确的负区间就停止更新。

代码修正建议

1. 加入特征归一化

对x_train做标准化处理,消除尺度差异,让梯度下降更容易收敛:

# 标准化特征
x_mean = np.mean(x_train)
x_std = np.std(x_train)
x_train_normalized = (x_train - x_mean) / x_std

后续将归一化后的x_train_normalized代入梯度下降函数。

2. 调整学习率与迭代次数

将学习率调大(如alpha=0.01),同时增加迭代次数(如num_iterations=10000),或添加早停机制监控损失变化:

# 修改参数
alpha = 0.01
num_iterations = 10000

3. 优化梯度下降函数(可选)

添加损失监控方便观察收敛过程,同时改用向量化计算提升效率:

def compute_gradient(x, y, w, b):
    m = len(y)
    predictions = w * x + b
    w_gradient = np.sum((predictions - y) * x) / m
    b_gradient = np.sum(predictions - y) / m
    return w_gradient, b_gradient

def gradient_descent(x, y, w, b, alpha, num_iterations):
    costs = []
    for i in range(num_iterations):
        w_gradient, b_gradient = compute_gradient(x, y, w, b)
        w -= alpha * w_gradient
        b -= alpha * b_gradient
        # 每100次迭代记录一次损失
        if i % 100 == 0:
            cost = compute_cost(x, y, w, b)
            costs.append(cost)
            print(f"Iteration {i}: Cost = {cost:.4f}")
    # 绘制损失曲线查看收敛情况
    plt.plot(costs)
    plt.xlabel("Iteration (x100)")
    plt.ylabel("Cost")
    plt.title("Cost vs Iterations")
    plt.show()
    return w, b

4. 修正绘图逻辑(可选)

原始绘图中plt.plot(x_train, w * x_train + b)会因x_train未排序导致折线混乱,建议先对x排序:

# 排序x和对应的预测值
sorted_indices = np.argsort(x_train)
x_sorted = x_train.iloc[sorted_indices]
y_pred_sorted = w * x_sorted + b

plt.scatter(x_train, y_train, marker='x', c='r', label='Data Points')
plt.plot(x_sorted, y_pred_sorted, '-', label='Linear Fit')

完整修正代码示例

import numpy as np
import matplotlib.pyplot as plt
from ISLP import load_data

# Load the Boston dataset
ds = load_data("Boston")
x_train, y_train = ds['lstat'], ds['medv']

# 标准化特征
x_mean = np.mean(x_train)
x_std = np.std(x_train)
x_train_normalized = (x_train - x_mean) / x_std

# Compute cost function
def compute_cost(x, y, w, b):
    m = len(y)
    predictions = w * x + b
    total_cost = np.sum((predictions - y) ** 2)
    return total_cost / (2 * m)

def compute_gradient(x, y, w, b):
    m = len(y)
    predictions = w * x + b
    w_gradient = np.sum((predictions - y) * x) / m
    b_gradient = np.sum(predictions - y) / m
    return w_gradient, b_gradient

# Gradient descent function with cost monitoring
def gradient_descent(x, y, w, b, alpha, num_iterations):
    costs = []
    for i in range(num_iterations):
        w_gradient, b_gradient = compute_gradient(x, y, w, b)
        w -= alpha * w_gradient
        b -= alpha * b_gradient
        if i % 100 == 0:
            cost = compute_cost(x, y, w, b)
            costs.append(cost)
            print(f"Iteration {i}: Cost = {cost:.4f}")
    # Plot cost curve
    plt.figure()
    plt.plot(costs)
    plt.xlabel("Iteration (x100)")
    plt.ylabel("Cost")
    plt.title("Cost vs Iterations")
    plt.show()
    return w, b

# Parameters
alpha = 0.01
num_iterations = 10000
initial_w = 0
initial_b = 0

# Run gradient descent on normalized data
w_norm, b_norm = gradient_descent(x_train_normalized, y_train, initial_w, initial_b, alpha, num_iterations)

# Convert normalized weight back to original scale
w = w_norm / x_std
b = b_norm - (w_norm * x_mean) / x_std

# Print the found parameters
print('w found by gradient descent:', w)
print('b found by gradient descent:', b)

# Plot with sorted x
sorted_indices = np.argsort(x_train)
x_sorted = x_train.iloc[sorted_indices]
y_pred_sorted = w * x_sorted + b

plt.scatter(x_train, y_train, marker='x', c='r', label='Data Points')
plt.plot(x_sorted, y_pred_sorted, '-', label='Linear Fit')
plt.title("Percentage of lower status people vs House Prices")
plt.ylabel("House Prices in thousands")
plt.xlabel("Percentage of lower status people")
plt.legend()
plt.show()

额外验证方法

可以用线性回归的解析解(最小二乘法)验证最优参数,确认梯度下降的收敛结果:

# 最小二乘法计算最优w和b
x_train_np = np.array(x_train).reshape(-1,1)
x_train_with_intercept = np.hstack([np.ones((x_train_np.shape[0],1)), x_train_np])
params_opt = np.linalg.inv(x_train_with_intercept.T @ x_train_with_intercept) @ x_train_with_intercept.T @ y_train
print(f"Optimal w (Least Squares): {params_opt[1]}")
print(f"Optimal b (Least Squares): {params_opt[0]}")

内容的提问来源于stack exchange,提问作者Dubem Nwokike

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 13:50:59