如何优化DataFrame行迭代效率?3万行数据处理耗时2小时求改进
3万行DataFrame数据处理提速方案
问题背景
现有3万行的DataFrame,需要遍历每行提取图片路径、统计特征和目标值,将图片读取缩放后存入数组。当前使用iterrows()遍历耗时约2小时,尝试apply()无法实现原有逻辑,寻求提速方案。原代码如下:
def get_X_y(df): X_pic = [] X_stats = [] y = [] new_size = (256, 256) # new_size=(width, height) counter = 0 #iterate over rows in dataframe for index,row in df.iterrows(): counter = counter + 1 print(counter) path = row['path'] img = cv2.imread(path) resize_img = cv2.resize(img, new_size) X_pic.append(resize_img) stats = row[['sex_1', 'sex_2', 'age_approx', 'site_1', 'site_2', 'site_3', 'site_4', 'site_5', 'site_6']] X_stats.append(stats) target = row['target'] y.append(target) X_pic, X_stats = np.array(X_pic).astype('float32'), np.array(X_stats).astype('float32') y = np.array(y).astype('float32') return (X_pic, X_stats), y # Get the training data (X_train_pic, X_train_stats), y_train = get_X_y(train_df) (X_train_pic.shape, X_train_stats.shape), y_train.shape # Get the val data (X_val_pic, X_val_stats), y_val = get_X_y(val_df) (X_val_pic.shape, X_val_stats.shape), y_val.shape
核心提速方案
1. 移除冗余打印操作
循环内的print(counter)会大幅拖慢处理速度,直接删除该代码块。
2. 批量提取统计特征与目标值
放弃逐行append的方式,直接利用DataFrame的向量化操作批量转换为numpy数组,速度提升显著:
# 替代原循环内X_stats和y的append逻辑 X_stats = df[['sex_1', 'sex_2', 'age_approx', 'site_1', 'site_2', 'site_3', 'site_4', 'site_5', 'site_6']].values.astype('float32') y = df['target'].values.astype('float32')
3. 多进程并行处理图片IO与缩放
图片读取和缩放是耗时瓶颈,用多进程并行处理可充分利用CPU多核资源:
先定义单张图片处理函数:
import cv2 import numpy as np from concurrent.futures import ProcessPoolExecutor def process_image(path, new_size): img = cv2.imread(path) if img is None: # 处理图片读取失败的情况,返回全0数组兜底 return np.zeros((new_size[1], new_size[0], 3), dtype='float32') resize_img = cv2.resize(img, new_size) return resize_img.astype('float32')
再修改get_X_y函数实现并行处理:
def get_X_y(df): new_size = (256, 256) # 批量处理统计特征和目标值 X_stats = df[['sex_1', 'sex_2', 'age_approx', 'site_1', 'site_2', 'site_3', 'site_4', 'site_5', 'site_6']].values.astype('float32') y = df['target'].values.astype('float32') # 多进程批量处理图片 with ProcessPoolExecutor() as executor: X_pic = list(executor.map(process_image, df['path'], [new_size]*len(df))) # 转换为numpy数组 X_pic = np.array(X_pic) return (X_pic, X_stats), y
4. 预分配numpy数组(可选优化)
若明确图片最终形状,可提前分配数组,避免列表append后的数组转换开销:
# 预分配数组替代list转array X_pic = np.zeros((len(df), new_size[1], new_size[0], 3), dtype='float32') with ProcessPoolExecutor() as executor: for idx, img in enumerate(executor.map(process_image, df['path'], [new_size]*len(df))): X_pic[idx] = img
效果预期
以上优化后,3万行数据的处理时间可压缩至十几分钟甚至更短,具体耗时取决于CPU核心数和磁盘IO性能。
内容的提问来源于stack exchange,提问作者bilbo_slagins
相关产品推荐
相关产品推荐

