You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

梯度下降异常:SVD降维特征集运行时长反而长于原特征集

SVD降维后梯度下降运行时长超出预期的问题分析

在Python实现梯度下降用于机器学习任务时,发现经SVD降维后的特征集,其梯度下降的运行时长反而比原特征集更长,调整学习率、增加迭代次数后该现象仍存在。以下是代码及排查分析:

代码实现

import time
import numpy as np
import matplotlib.pyplot as plt

def h_theta(X1, theta1):
    # 假设函数实现
    return np.dot(X1, theta1)

def j_theta(X1, y1, theta1):
    # 代价函数实现
    return np.sum((h_theta(X1, theta1) - y1) ** 2) / (2 * X1.size)

def grad(X1, y1, theta):
    # 梯度计算
    gradient = np.dot(X1.T, h_theta(X1, theta) - y1) / len(y1)
    return gradient

def gradient_descent(X1, y1):
    theta_initial = np.zeros(X1.shape[1])  # 初始化theta为0

    num_iterations = 1000
    learning_rates = [0.1, 0.01, 0.001]  
    cost_iterations = []
    theta_values = []
    start = time.time()
    for alpha in learning_rates:
        theta = theta_initial.copy()
        cost_history = []
        for i in range(num_iterations):
            gradient = grad(X1, y1, theta)
            theta = theta - np.dot(alpha, gradient)
            cost = j_theta(X1, y1, theta)
            cost_history.append(cost)
        cost_iterations.append(cost_history)
        theta_values.append(theta)
    end = time.time()
    print(f"Time taken: {end - start} seconds")
    fig, axs = plt.subplots(len(learning_rates), figsize=(8, 15))  
    for i, alpha in enumerate(learning_rates):
        axs[i].plot(range(num_iterations), cost_iterations[i], label=f'alpha = {alpha}')  
        axs[i].set_title(f'Learning Rate: {alpha}')
        axs[i].set_ylabel('Cost J')
        axs[i].set_xlabel('Number of Iterations')  
        axs[i].legend()
    plt.tight_layout()  
    plt.show()

# SVD降维至3个特征
U, S, Vt = np.linalg.svd(X_normalized)
X_reduced = np.dot(X_normalized, Vt[:3].T)

print("First 5 rows of X_reduced:")
# 归一化降维后的数据
X_reduced = (X_reduced - np.mean(X_reduced, axis=0)) / np.std(X_reduced, axis=0)

print("the means and stds of X after being reduced and normalized:\n" ,X_reduced.mean(axis=0), X_reduced.std(axis=0))
print("Shape of X_reduced:", X_reduced.shape)

# 添加截距列
X_reduced_with_intercept = np.hstack((intercept_column, X_reduced))

# 对原数据集和降维数据集分别执行梯度下降
gradient_descent(X_normalized_with_intercept, y_normalized)
gradient_descent(X_reduced_with_intercept, y_normalized)

可能的原因分析

  1. 代价函数分母逻辑异常:代价函数j_theta使用X1.size(特征矩阵总元素数)作为分母,而非标准均方误差的样本数len(y1)。降维后X1.size更小,可能导致代价数值波动更大,间接影响numpy内部运算优化效率。
  2. 内存布局与数据类型差异:SVD降维后的矩阵可能与原数据集内存布局(如C连续/Fortran连续)不一致,numpy对不同布局数组的矩阵乘法效率有差异;若降维过程中发生隐性数据类型转换,也会增加运算开销。
  3. 缓存命中差异:首次运行原数据集时,部分数据已加载到CPU缓存,而降维后的新数据集未命中缓存,导致运算延迟增加。
  4. 梯度计算的数值特性:SVD降维后的特征可能存在隐性共线性或数值异常(尽管做了归一化),使梯度计算的浮点运算复杂度提升。

排查建议

  • 精准计时核心运算:将计时范围缩小到梯度下降迭代循环内部,排除学习率循环、绘图等非核心代码干扰,单独统计每次迭代的运算时间。
  • 检查数据属性:用X_normalized_with_intercept.dtype、X_reduced_with_intercept.dtype确认数据类型一致;用np.iscontiguous()检查内存布局,必要时用np.ascontiguousarray()转换为C连续布局。
  • 简化测试场景:暂时移除绘图代码,仅保留梯度计算与计时逻辑;使用小样本数据集测试,观察是否仍存在降维后更慢的现象。
  • 修正代价函数:将代价函数分母改为样本数len(y1),符合标准MSE定义,消除分母差异带来的潜在运算影响。
  • 定位耗时环节:单独对比原数据集与降维数据集的矩阵乘法(np.dot(X, theta))、转置点积(np.dot(X.T, error))运算时间,找到具体耗时步骤。

内容的提问来源于stack exchange,提问作者H.S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 15:33:16