使用Python Pandas时,向量化梯度下降为何比循环版慢很多?
向量化版本代码
import pandas as pd import numpy as np alpha=.001 data = [ (2, 3), (4, 7), (6, 11), (8, 17), (10, 23), (12, 31), (14, 39), (16, 49), (18, 59), (20, 71), (22, 83), (24, 97), (26, 113), (28, 131), (30, 149), (32, 169), (34, 191), (36, 214), (38, 239), (40, 266), (42, 295), (44, 326), (46, 359), (48, 394), (50, 431)] ##CREATE EXAMPLES MATRIX x_coordinates = [x[0] for x in data] x_coords=[] [x_coords.append([1,x]) for x in x_coordinates] #Creates a list of all x-coordinates with a 1 column examples=pd.DataFrame(x_coords).transpose() #uses that list to create a dataframe. Must transpose so it is dimsion 2,25. rows are features, columns are specific examples. ##CREATE THETA MATRIX/VECTOR theta_list = [1, 2] theta = pd.DataFrame(theta_list).transpose() #creates a df of dimension 1,2. ##CREATE Y VECTOR/MATRIX y_coordinates = [x[1] for x in data] y=pd.DataFrame(y_coordinates).transpose() deriv=pd.DataFrame([]) count=0 while (deriv != 0).all().all() and count <= 500000: length=len(data) #theta*X thetaX=theta.dot(examples) error=thetaX-y error_pt2=error.dot(examples.T) deriv=alpha*(1/length)*error_pt2 theta=theta-deriv print(theta) count+=1 print(count)
循环版本代码
total=0 th0=0 th1=0 alpha=0.001 deriv0=1 deriv1=1 count=0 while deriv0 and deriv1 != 0 and count<=1000000: total0=0 total1=0 #th0 for i in data: hyp=th0+(th1*i[0]) #print("Hyp is {}".format(hyp)) total0+=(hyp-i[1]) deriv0=(1/25)*total0 th0temp=th0-(alpha*(deriv0)) #th1 for i in data: hyp=th0+(th1*i[0]) total1+=(hyp-i[1])*i[0] deriv1=(1/25)*total1 th1temp=th1-(alpha*(deriv1)) th0=th0temp th1=th1temp th0temp=0 th1temp=0 count+=1 print("Theta 0: {} \n Theta 1: {} \n\n".format(th0,th1)) print(count)
疑问
我运行向量化版本时,耗时几乎是循环版的10倍。我本以为向量化实现会比多循环版本高效得多,这是为什么?是否是Pandas的计算开销导致速度变慢?或许Pandas并不适合这类算法?
问题分析与优化方案
为什么向量化版本更慢?
核心原因是Pandas DataFrame的额外开销:DataFrame本质是带索引、列名等元数据的表格结构,专为结构化数据处理设计,而非底层数值计算。每次矩阵运算时,Pandas需要维护这些元数据,还要处理DataFrame的内部逻辑,这些开销在小数据集(仅25条样本)下,完全盖过了向量化运算的效率优势。
而你的循环版本用的是原生Python变量和列表,数据结构极度轻量化,没有额外的维护成本,在小样本量下反而跑得更快。
Pandas是否适合这类算法?
Pandas并不适合纯数值矩阵运算的场景,这类任务应该交给NumPy——它是专为数值计算优化的库,数组结构没有冗余元数据,向量化运算直接调用底层C实现,效率极高。你的向量化版本选错了工具,才导致速度不如循环版。
优化后的向量化版本示例
把所有DataFrame替换为NumPy数组,修改后的代码如下:
import numpy as np alpha = 0.001 data = [ (2, 3), (4, 7), (6, 11), (8, 17), (10, 23), (12, 31), (14, 39), (16, 49), (18, 59), (20, 71), (22, 83), (24, 97), (26, 113), (28, 131), (30, 149), (32, 169), (34, 191), (36, 214), (38, 239), (40, 266), (42, 295), (44, 326), (46, 359), (48, 394), (50, 431) ] # 创建特征矩阵(2行25列) x_coordinates = [x[0] for x in data] x_coords = [[1, x] for x in x_coordinates] examples = np.array(x_coords).T # 创建theta向量(1行2列) theta = np.array([[1, 2]]) # 创建y向量(1行25列) y_coordinates = [x[1] for x in data] y = np.array([y_coordinates]) deriv = np.array([]) count = 0 length = len(data) while (deriv != 0).all() if deriv.size > 0 else True and count <= 500000: thetaX = theta.dot(examples) error = thetaX - y error_pt2 = error.dot(examples.T) deriv = alpha * (1 / length) * error_pt2 theta = theta - deriv print(theta) count += 1 print(count)
这个版本用NumPy替代Pandas后,向量化运算的效率会远超你的循环版本,尤其是当样本量增大时,优势会更明显。
额外说明
向量化的优势在于大数据量场景:当样本量达到数千、上万级别时,Python循环的开销会急剧上升,而NumPy的向量化运算能保持线性增长的效率。小样本量下,轻量化的循环确实可能更快,但这不是向量化的问题,而是工具选择和场景匹配的问题。
内容的提问来源于stack exchange,提问作者Jcb Rb
相关产品推荐
相关产品推荐

