线性回归手动梯度下降处理大数值时遇标量幂溢出问题求助
问题场景
手动实现线性回归梯度下降处理房屋价格数据集(X=LT,Y=Harga)时,出现overflow encountered in scalar power等数值溢出警告,具体报错如下:
RuntimeWarning: overflow encountered in scalar add
cost = np.float64(cost + (f_wb - y[i])**2)
RuntimeWarning: overflow encountered in scalar power
cost = np.float64(cost + (f_wb - y[i])**2)
RuntimeWarning: overflow encountered in scalar add
dj_dw += np.float64(dj_dw_i)
RuntimeWarning: invalid value encountered in scalar subtract
w = np.float64(w - alpha * dj_dw)
数据导入代码:
import os import openpyxl from openpyxl import Workbook import numpy as np wb = openpyxl.load_workbook('DATA RUMAH.xlsx') ws = wb.active y_train_data = np.array([]) x_train_data = np.array([]) def get_x_train(): x_train = np.array([]) # Initialize x_train as a local variable for x in range(2, 1011): data = ws.cell(row=x, column=5).value x_train = np.append(x_train, data) return x_train def get_y_train(): y_train = np.array([]) # Initialize y_train as a local variable for y in range(2, 1011): data = ws.cell(row=y, column=3).value y_train = np.append(y_train, data) return y_train
完整梯度下降实现代码:
import math, copy import numpy as np import matplotlib.pyplot as plt plt.style.use('./deeplearning.mplstyle') from lab_utils_uni import plt_house_x, plt_contour_wgrad, plt_divergence, plt_gradients # Load our data set x_train = get_x_train() #features y_train = get_y_train() #target value #Function to calculate the cost def compute_cost(x, y, w, b): m = x.shape[0] cost = 0 for i in range(m): f_wb = w * x[i] + b cost = cost + (f_wb - y[i])**2 total_cost = 1 / (2 * m) * cost return total_cost def compute_gradient(x, y, w, b): """ Computes the gradient for linear regression Args: x (ndarray (m,)): Data, m examples y (ndarray (m,)): target values w,b (scalar) : model parameters Returns dj_dw (scalar): The gradient of the cost w.r.t. the parameters w dj_db (scalar): The gradient of the cost w.r.t. the parameter b """ # Number of training examples m = x.shape[0] dj_dw = 0 dj_db = 0 for i in range(m): f_wb = w * x[i] + b dj_dw_i = (f_wb - y[i]) * x[i] dj_db_i = f_wb - y[i] dj_db += dj_db_i dj_dw += dj_dw_i dj_dw = dj_dw / m dj_db = dj_db / m return dj_dw, dj_db def gradient_descent(x, y, w_in, b_in, alpha, num_iters, cost_function, gradient_function): """ Performs gradient descent to fit w,b. Updates w,b by taking num_iters gradient steps with learning rate alpha Args: x (ndarray (m,)) : Data, m examples y (ndarray (m,)) : target values w_in,b_in (scalar): initial values of model parameters alpha (float): Learning rate num_iters (int): number of iterations to run gradient descent cost_function: function to call to produce cost gradient_function: function to call to produce gradient Returns: w (scalar): Updated value of parameter after running gradient descent b (scalar): Updated value of parameter after running gradient descent J_history (List): History of cost values p_history (list): History of parameters [w,b] """ # Specify data type as np.float64 for w, b w = np.float64(w_in) b = np.float64(b_in) # An array to store cost J and w's at each iteration primarily for graphing later J_history = [] p_history = [] for i in range(num_iters): # Calculate the gradient and update the parameters using gradient_function dj_dw, dj_db = gradient_function(x, y, w , b) # Update Parameters using equation (3) above b = b - alpha * dj_db w = w - alpha * dj_dw # Save cost J at each iteration J_history.append(cost_function(x, y, w, b)) p_history.append([w, b]) # Print cost every at intervals 10 times or as many iterations if < 10 if i % math.ceil(num_iters/10) == 0: print(f"Iteration {i:4}: Cost {J_history[-1]:0.2e} ", f"dj_dw: {dj_dw: 0.3e}, dj_db: {dj_db: 0.3e} ", f"w: {w: 0.3e}, b: {b: 0.5e}") return w, b, J_history, p_history # Return w and J,w history for graphing # Initialize parameters with np.float64 data type w_init = np.float64(0) b_init = np.float64(0) # Some gradient descent settings iterations = 100000 tmp_alpha = np.float64(1.0e-4) # Run gradient descent w_final, b_final, J_hist, p_hist = gradient_descent(x_train, y_train, w_init, b_init, tmp_alpha, iterations, compute_cost, compute_gradient) # Print the result print(f"(w, b) found by gradient descent: ({w_final:8.4f}, {b_final:8.4f})")
用户曾尝试归一化但担心过度扭曲数据,希望找到可行的大数值数据处理方案。
解决方案
1. 正确使用特征缩放(不会过度扭曲数据)
特征缩放是处理大数值数据的标准手段,不会破坏数据的相对关系,反而能让梯度下降稳定收敛。推荐两种常用方式:
- 标准化(Z-Score):将数据转换为均值0、标准差1的分布,代码实现:
训练完成后,只需反缩放即可还原真实预测值:# 缩放特征和目标值 x_mean = np.mean(x_train) x_std = np.std(x_train) x_scaled = (x_train - x_mean) / x_std y_mean = np.mean(y_train) y_std = np.std(y_train) y_scaled = (y_train - y_mean) / y_std# 假设用缩放后的数据得到w_scaled, b_scaled y_pred_scaled = w_scaled * x_scaled + b_scaled y_pred = y_pred_scaled * y_std + y_mean - 最小-最大归一化:将数据缩放到[0,1]区间,适合数据分布明确的场景:
反缩放逻辑类似:x_min = np.min(x_train) x_max = np.max(x_train) x_scaled = (x_train - x_min) / (x_max - x_min) y_min = np.min(y_train) y_max = np.max(y_train) y_scaled = (y_train - y_min) / (y_max - y_min)y_pred = y_pred_scaled * (y_max - y_min) + y_min
2. 调整学习率
当前设置的1.0e-4对于大数值数据仍然偏大,会导致参数更新幅度过大引发溢出。可以尝试逐步降低学习率,比如1.0e-7、1.0e-8,同时减少迭代次数(如10000次)观察收敛情况。
3. 用向量化运算替代循环
原代码中的循环计算效率低且容易累积数值误差,改用numpy向量化操作可提升效率并减少数值问题:
def compute_cost(x, y, w, b): m = x.shape[0] f_wb = w * x + b cost = np.sum((f_wb - y)**2) total_cost = cost / (2 * m) return total_cost def compute_gradient(x, y, w, b): m = x.shape[0] f_wb = w * x + b dj_dw = np.sum((f_wb - y) * x) / m dj_db = np.sum(f_wb - y) / m return dj_dw, dj_db
4. 监测收敛情况
在梯度下降过程中,观察成本函数J_history的变化:
- 若成本持续下降,说明模型在收敛;
- 若成本波动或上升,需降低学习率;
- 若成本下降极慢,可适当调升学习率。
添加代码绘制成本变化曲线辅助判断:
plt.plot(J_hist) plt.xlabel("迭代次数") plt.ylabel("成本值") plt.title("成本随迭代次数变化") plt.show()
内容的提问来源于stack exchange,提问作者Fauzan Anggito Wicakson

