梯度下降异常:SVD降维特征集运行时长反而长于原特征集
SVD降维后梯度下降运行时长超出预期的问题分析
在Python实现梯度下降用于机器学习任务时,发现经SVD降维后的特征集,其梯度下降的运行时长反而比原特征集更长,调整学习率、增加迭代次数后该现象仍存在。以下是代码及排查分析:
代码实现
import time import numpy as np import matplotlib.pyplot as plt def h_theta(X1, theta1): # 假设函数实现 return np.dot(X1, theta1) def j_theta(X1, y1, theta1): # 代价函数实现 return np.sum((h_theta(X1, theta1) - y1) ** 2) / (2 * X1.size) def grad(X1, y1, theta): # 梯度计算 gradient = np.dot(X1.T, h_theta(X1, theta) - y1) / len(y1) return gradient def gradient_descent(X1, y1): theta_initial = np.zeros(X1.shape[1]) # 初始化theta为0 num_iterations = 1000 learning_rates = [0.1, 0.01, 0.001] cost_iterations = [] theta_values = [] start = time.time() for alpha in learning_rates: theta = theta_initial.copy() cost_history = [] for i in range(num_iterations): gradient = grad(X1, y1, theta) theta = theta - np.dot(alpha, gradient) cost = j_theta(X1, y1, theta) cost_history.append(cost) cost_iterations.append(cost_history) theta_values.append(theta) end = time.time() print(f"Time taken: {end - start} seconds") fig, axs = plt.subplots(len(learning_rates), figsize=(8, 15)) for i, alpha in enumerate(learning_rates): axs[i].plot(range(num_iterations), cost_iterations[i], label=f'alpha = {alpha}') axs[i].set_title(f'Learning Rate: {alpha}') axs[i].set_ylabel('Cost J') axs[i].set_xlabel('Number of Iterations') axs[i].legend() plt.tight_layout() plt.show() # SVD降维至3个特征 U, S, Vt = np.linalg.svd(X_normalized) X_reduced = np.dot(X_normalized, Vt[:3].T) print("First 5 rows of X_reduced:") # 归一化降维后的数据 X_reduced = (X_reduced - np.mean(X_reduced, axis=0)) / np.std(X_reduced, axis=0) print("the means and stds of X after being reduced and normalized:\n" ,X_reduced.mean(axis=0), X_reduced.std(axis=0)) print("Shape of X_reduced:", X_reduced.shape) # 添加截距列 X_reduced_with_intercept = np.hstack((intercept_column, X_reduced)) # 对原数据集和降维数据集分别执行梯度下降 gradient_descent(X_normalized_with_intercept, y_normalized) gradient_descent(X_reduced_with_intercept, y_normalized)
可能的原因分析
- 代价函数分母逻辑异常:代价函数
j_theta使用X1.size(特征矩阵总元素数)作为分母,而非标准均方误差的样本数len(y1)。降维后X1.size更小,可能导致代价数值波动更大,间接影响numpy内部运算优化效率。 - 内存布局与数据类型差异:SVD降维后的矩阵可能与原数据集内存布局(如C连续/Fortran连续)不一致,numpy对不同布局数组的矩阵乘法效率有差异;若降维过程中发生隐性数据类型转换,也会增加运算开销。
- 缓存命中差异:首次运行原数据集时,部分数据已加载到CPU缓存,而降维后的新数据集未命中缓存,导致运算延迟增加。
- 梯度计算的数值特性:SVD降维后的特征可能存在隐性共线性或数值异常(尽管做了归一化),使梯度计算的浮点运算复杂度提升。
排查建议
- 精准计时核心运算:将计时范围缩小到梯度下降迭代循环内部,排除学习率循环、绘图等非核心代码干扰,单独统计每次迭代的运算时间。
- 检查数据属性:用
X_normalized_with_intercept.dtype、X_reduced_with_intercept.dtype确认数据类型一致;用np.iscontiguous()检查内存布局,必要时用np.ascontiguousarray()转换为C连续布局。 - 简化测试场景:暂时移除绘图代码,仅保留梯度计算与计时逻辑;使用小样本数据集测试,观察是否仍存在降维后更慢的现象。
- 修正代价函数:将代价函数分母改为样本数
len(y1),符合标准MSE定义,消除分母差异带来的潜在运算影响。 - 定位耗时环节:单独对比原数据集与降维数据集的矩阵乘法(
np.dot(X, theta))、转置点积(np.dot(X.T, error))运算时间,找到具体耗时步骤。
内容的提问来源于stack exchange,提问作者H.S
相关产品推荐
相关产品推荐

