You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

线性回归手动梯度下降处理大数值时遇标量幂溢出问题求助

梯度下降处理大数值数据时的溢出问题解决

问题场景

手动实现线性回归梯度下降处理房屋价格数据集(X=LT,Y=Harga)时,出现overflow encountered in scalar power等数值溢出警告,具体报错如下:

RuntimeWarning: overflow encountered in scalar add
cost = np.float64(cost + (f_wb - y[i])**2)
RuntimeWarning: overflow encountered in scalar power
cost = np.float64(cost + (f_wb - y[i])**2)
RuntimeWarning: overflow encountered in scalar add
dj_dw += np.float64(dj_dw_i)
RuntimeWarning: invalid value encountered in scalar subtract
w = np.float64(w - alpha * dj_dw)

数据导入代码:

import os
import openpyxl
from openpyxl import Workbook
import numpy as np

wb = openpyxl.load_workbook('DATA RUMAH.xlsx')
ws = wb.active

y_train_data = np.array([])
x_train_data = np.array([])

def get_x_train():
    x_train = np.array([])  # Initialize x_train as a local variable
    for x in range(2, 1011):
        data = ws.cell(row=x, column=5).value
        x_train = np.append(x_train, data)
    return x_train

def get_y_train():
    y_train = np.array([])  # Initialize y_train as a local variable
    for y in range(2, 1011):
        data = ws.cell(row=y, column=3).value
        y_train = np.append(y_train, data)
    return y_train

完整梯度下降实现代码:

import math, copy
import numpy as np
import matplotlib.pyplot as plt
plt.style.use('./deeplearning.mplstyle')
from lab_utils_uni import plt_house_x, plt_contour_wgrad, plt_divergence, plt_gradients

# Load our data set
x_train = get_x_train()   #features
y_train = get_y_train()  #target value

#Function to calculate the cost
def compute_cost(x, y, w, b):
   
    m = x.shape[0] 
    cost = 0
    
    for i in range(m):
        f_wb = w * x[i] + b
        cost = cost + (f_wb - y[i])**2
    total_cost = 1 / (2 * m) * cost

    return total_cost

def compute_gradient(x, y, w, b): 
    """
    Computes the gradient for linear regression 
    Args:
      x (ndarray (m,)): Data, m examples 
      y (ndarray (m,)): target values
      w,b (scalar)    : model parameters  
    Returns
      dj_dw (scalar): The gradient of the cost w.r.t. the parameters w
      dj_db (scalar): The gradient of the cost w.r.t. the parameter b     
     """
    
    # Number of training examples
    m = x.shape[0]    
    dj_dw = 0
    dj_db = 0
    
    for i in range(m):  
        f_wb = w * x[i] + b 
        dj_dw_i = (f_wb - y[i]) * x[i] 
        dj_db_i = f_wb - y[i] 
        dj_db += dj_db_i
        dj_dw += dj_dw_i 
    dj_dw = dj_dw / m 
    dj_db = dj_db / m 
        
    return dj_dw, dj_db

def gradient_descent(x, y, w_in, b_in, alpha, num_iters, cost_function, gradient_function):
    """
    Performs gradient descent to fit w,b. Updates w,b by taking 
    num_iters gradient steps with learning rate alpha
    
    Args:
        x (ndarray (m,))  : Data, m examples 
        y (ndarray (m,))  : target values
        w_in,b_in (scalar): initial values of model parameters  
        alpha (float):     Learning rate
        num_iters (int):   number of iterations to run gradient descent
        cost_function:     function to call to produce cost
        gradient_function: function to call to produce gradient
      
    Returns:
        w (scalar): Updated value of parameter after running gradient descent
        b (scalar): Updated value of parameter after running gradient descent
        J_history (List): History of cost values
        p_history (list): History of parameters [w,b] 
    """
    
    # Specify data type as np.float64 for w, b
    w = np.float64(w_in)
    b = np.float64(b_in)
    
    # An array to store cost J and w's at each iteration primarily for graphing later
    J_history = []
    p_history = []
    
    for i in range(num_iters):
        # Calculate the gradient and update the parameters using gradient_function
        dj_dw, dj_db = gradient_function(x, y, w , b)     

        # Update Parameters using equation (3) above
        b = b - alpha * dj_db                            
        w = w - alpha * dj_dw                            

        # Save cost J at each iteration
        J_history.append(cost_function(x, y, w, b))
        p_history.append([w, b])

        # Print cost every at intervals 10 times or as many iterations if < 10
        if i % math.ceil(num_iters/10) == 0:
            print(f"Iteration {i:4}: Cost {J_history[-1]:0.2e} ",
                  f"dj_dw: {dj_dw: 0.3e}, dj_db: {dj_db: 0.3e}  ",
                  f"w: {w: 0.3e}, b: {b: 0.5e}")
 
    return w, b, J_history, p_history  # Return w and J,w history for graphing

# Initialize parameters with np.float64 data type
w_init = np.float64(0)
b_init = np.float64(0)

# Some gradient descent settings
iterations = 100000
tmp_alpha = np.float64(1.0e-4)

# Run gradient descent
w_final, b_final, J_hist, p_hist = gradient_descent(x_train, y_train, w_init, b_init, tmp_alpha,
                                                    iterations, compute_cost, compute_gradient)

# Print the result
print(f"(w, b) found by gradient descent: ({w_final:8.4f}, {b_final:8.4f})")

用户曾尝试归一化但担心过度扭曲数据,希望找到可行的大数值数据处理方案。


解决方案

1. 正确使用特征缩放(不会过度扭曲数据)

特征缩放是处理大数值数据的标准手段,不会破坏数据的相对关系,反而能让梯度下降稳定收敛。推荐两种常用方式:

  • 标准化(Z-Score):将数据转换为均值0、标准差1的分布,代码实现:
    # 缩放特征和目标值
    x_mean = np.mean(x_train)
    x_std = np.std(x_train)
    x_scaled = (x_train - x_mean) / x_std
    
    y_mean = np.mean(y_train)
    y_std = np.std(y_train)
    y_scaled = (y_train - y_mean) / y_std
    
    训练完成后,只需反缩放即可还原真实预测值:
    # 假设用缩放后的数据得到w_scaled, b_scaled
    y_pred_scaled = w_scaled * x_scaled + b_scaled
    y_pred = y_pred_scaled * y_std + y_mean
    
  • 最小-最大归一化:将数据缩放到[0,1]区间,适合数据分布明确的场景:
    x_min = np.min(x_train)
    x_max = np.max(x_train)
    x_scaled = (x_train - x_min) / (x_max - x_min)
    
    y_min = np.min(y_train)
    y_max = np.max(y_train)
    y_scaled = (y_train - y_min) / (y_max - y_min)
    
    反缩放逻辑类似:
    y_pred = y_pred_scaled * (y_max - y_min) + y_min
    

2. 调整学习率

当前设置的1.0e-4对于大数值数据仍然偏大,会导致参数更新幅度过大引发溢出。可以尝试逐步降低学习率,比如1.0e-7、1.0e-8,同时减少迭代次数(如10000次)观察收敛情况。

3. 用向量化运算替代循环

原代码中的循环计算效率低且容易累积数值误差,改用numpy向量化操作可提升效率并减少数值问题:

def compute_cost(x, y, w, b):
    m = x.shape[0]
    f_wb = w * x + b
    cost = np.sum((f_wb - y)**2)
    total_cost = cost / (2 * m)
    return total_cost

def compute_gradient(x, y, w, b):
    m = x.shape[0]
    f_wb = w * x + b
    dj_dw = np.sum((f_wb - y) * x) / m
    dj_db = np.sum(f_wb - y) / m
    return dj_dw, dj_db

4. 监测收敛情况

在梯度下降过程中,观察成本函数J_history的变化:

  • 若成本持续下降,说明模型在收敛;
  • 若成本波动或上升,需降低学习率;
  • 若成本下降极慢,可适当调升学习率。

添加代码绘制成本变化曲线辅助判断:

plt.plot(J_hist)
plt.xlabel("迭代次数")
plt.ylabel("成本值")
plt.title("成本随迭代次数变化")
plt.show()

内容的提问来源于stack exchange,提问作者Fauzan Anggito Wicakson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 14:37:09