Gradient Descent拟合异常:散点负相关但拟合线正相关求助
梯度下降拟合线与数据趋势相反的问题排查与解决
问题根源分析
你的拟合线呈现正相关但数据实际为负相关,核心原因是特征未做归一化+学习率过小+迭代次数不足,导致梯度下降在有限迭代内无法收敛到正确参数:
- Boston数据集的
lstat(低收入人群占比)范围约0-37,medv(房价)范围约5-50,两者尺度差异极大,未归一化时梯度量级会异常大,小学习率下参数更新极慢。 - 初始参数
w=0时,梯度计算结果为负(prediction - y为负,乘以正的x后求和平均为负),按梯度下降规则w -= alpha * w_gradient,会让w往正方向小幅增加,但真实最优w是负数,小学习率+少迭代次数下,参数还未走到正确的负区间就停止更新。
代码修正建议
1. 加入特征归一化
对x_train做标准化处理,消除尺度差异,让梯度下降更容易收敛:
# 标准化特征 x_mean = np.mean(x_train) x_std = np.std(x_train) x_train_normalized = (x_train - x_mean) / x_std
后续将归一化后的x_train_normalized代入梯度下降函数。
2. 调整学习率与迭代次数
将学习率调大(如alpha=0.01),同时增加迭代次数(如num_iterations=10000),或添加早停机制监控损失变化:
# 修改参数 alpha = 0.01 num_iterations = 10000
3. 优化梯度下降函数(可选)
添加损失监控方便观察收敛过程,同时改用向量化计算提升效率:
def compute_gradient(x, y, w, b): m = len(y) predictions = w * x + b w_gradient = np.sum((predictions - y) * x) / m b_gradient = np.sum(predictions - y) / m return w_gradient, b_gradient def gradient_descent(x, y, w, b, alpha, num_iterations): costs = [] for i in range(num_iterations): w_gradient, b_gradient = compute_gradient(x, y, w, b) w -= alpha * w_gradient b -= alpha * b_gradient # 每100次迭代记录一次损失 if i % 100 == 0: cost = compute_cost(x, y, w, b) costs.append(cost) print(f"Iteration {i}: Cost = {cost:.4f}") # 绘制损失曲线查看收敛情况 plt.plot(costs) plt.xlabel("Iteration (x100)") plt.ylabel("Cost") plt.title("Cost vs Iterations") plt.show() return w, b
4. 修正绘图逻辑(可选)
原始绘图中plt.plot(x_train, w * x_train + b)会因x_train未排序导致折线混乱,建议先对x排序:
# 排序x和对应的预测值 sorted_indices = np.argsort(x_train) x_sorted = x_train.iloc[sorted_indices] y_pred_sorted = w * x_sorted + b plt.scatter(x_train, y_train, marker='x', c='r', label='Data Points') plt.plot(x_sorted, y_pred_sorted, '-', label='Linear Fit')
完整修正代码示例
import numpy as np import matplotlib.pyplot as plt from ISLP import load_data # Load the Boston dataset ds = load_data("Boston") x_train, y_train = ds['lstat'], ds['medv'] # 标准化特征 x_mean = np.mean(x_train) x_std = np.std(x_train) x_train_normalized = (x_train - x_mean) / x_std # Compute cost function def compute_cost(x, y, w, b): m = len(y) predictions = w * x + b total_cost = np.sum((predictions - y) ** 2) return total_cost / (2 * m) def compute_gradient(x, y, w, b): m = len(y) predictions = w * x + b w_gradient = np.sum((predictions - y) * x) / m b_gradient = np.sum(predictions - y) / m return w_gradient, b_gradient # Gradient descent function with cost monitoring def gradient_descent(x, y, w, b, alpha, num_iterations): costs = [] for i in range(num_iterations): w_gradient, b_gradient = compute_gradient(x, y, w, b) w -= alpha * w_gradient b -= alpha * b_gradient if i % 100 == 0: cost = compute_cost(x, y, w, b) costs.append(cost) print(f"Iteration {i}: Cost = {cost:.4f}") # Plot cost curve plt.figure() plt.plot(costs) plt.xlabel("Iteration (x100)") plt.ylabel("Cost") plt.title("Cost vs Iterations") plt.show() return w, b # Parameters alpha = 0.01 num_iterations = 10000 initial_w = 0 initial_b = 0 # Run gradient descent on normalized data w_norm, b_norm = gradient_descent(x_train_normalized, y_train, initial_w, initial_b, alpha, num_iterations) # Convert normalized weight back to original scale w = w_norm / x_std b = b_norm - (w_norm * x_mean) / x_std # Print the found parameters print('w found by gradient descent:', w) print('b found by gradient descent:', b) # Plot with sorted x sorted_indices = np.argsort(x_train) x_sorted = x_train.iloc[sorted_indices] y_pred_sorted = w * x_sorted + b plt.scatter(x_train, y_train, marker='x', c='r', label='Data Points') plt.plot(x_sorted, y_pred_sorted, '-', label='Linear Fit') plt.title("Percentage of lower status people vs House Prices") plt.ylabel("House Prices in thousands") plt.xlabel("Percentage of lower status people") plt.legend() plt.show()
额外验证方法
可以用线性回归的解析解(最小二乘法)验证最优参数,确认梯度下降的收敛结果:
# 最小二乘法计算最优w和b x_train_np = np.array(x_train).reshape(-1,1) x_train_with_intercept = np.hstack([np.ones((x_train_np.shape[0],1)), x_train_np]) params_opt = np.linalg.inv(x_train_with_intercept.T @ x_train_with_intercept) @ x_train_with_intercept.T @ y_train print(f"Optimal w (Least Squares): {params_opt[1]}") print(f"Optimal b (Least Squares): {params_opt[0]}")
内容的提问来源于stack exchange,提问作者Dubem Nwokike
相关产品推荐
相关产品推荐

