You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化DataFrame行迭代效率?3万行数据处理耗时2小时求改进

3万行DataFrame数据处理提速方案

问题背景

现有3万行的DataFrame,需要遍历每行提取图片路径、统计特征和目标值,将图片读取缩放后存入数组。当前使用iterrows()遍历耗时约2小时,尝试apply()无法实现原有逻辑,寻求提速方案。原代码如下:

def get_X_y(df):

    X_pic = []
    X_stats = []
    y = []

    new_size = (256, 256) # new_size=(width, height)

    counter = 0
    #iterate over rows in dataframe

    for index,row in df.iterrows():

        counter = counter + 1
        print(counter)

        path = row['path']
        img = cv2.imread(path)
        resize_img = cv2.resize(img, new_size)
        X_pic.append(resize_img)

        stats = row[['sex_1', 'sex_2', 'age_approx', 'site_1', 'site_2', 'site_3', 'site_4', 'site_5', 'site_6']]
        X_stats.append(stats)
        
        target = row['target']
        y.append(target)

    X_pic, X_stats = np.array(X_pic).astype('float32'), np.array(X_stats).astype('float32')
    y = np.array(y).astype('float32')

    return (X_pic, X_stats), y

# Get the training data
(X_train_pic, X_train_stats), y_train = get_X_y(train_df)
(X_train_pic.shape, X_train_stats.shape), y_train.shape

# Get the val data
(X_val_pic, X_val_stats), y_val = get_X_y(val_df)
(X_val_pic.shape, X_val_stats.shape), y_val.shape

核心提速方案

1. 移除冗余打印操作

循环内的print(counter)会大幅拖慢处理速度,直接删除该代码块。

2. 批量提取统计特征与目标值

放弃逐行append的方式,直接利用DataFrame的向量化操作批量转换为numpy数组,速度提升显著:

# 替代原循环内X_stats和y的append逻辑
X_stats = df[['sex_1', 'sex_2', 'age_approx', 'site_1', 'site_2', 'site_3', 'site_4', 'site_5', 'site_6']].values.astype('float32')
y = df['target'].values.astype('float32')

3. 多进程并行处理图片IO与缩放

图片读取和缩放是耗时瓶颈,用多进程并行处理可充分利用CPU多核资源:
先定义单张图片处理函数:

import cv2
import numpy as np
from concurrent.futures import ProcessPoolExecutor

def process_image(path, new_size):
    img = cv2.imread(path)
    if img is None:
        # 处理图片读取失败的情况,返回全0数组兜底
        return np.zeros((new_size[1], new_size[0], 3), dtype='float32')
    resize_img = cv2.resize(img, new_size)
    return resize_img.astype('float32')

再修改get_X_y函数实现并行处理:

def get_X_y(df):
    new_size = (256, 256)
    
    # 批量处理统计特征和目标值
    X_stats = df[['sex_1', 'sex_2', 'age_approx', 'site_1', 'site_2', 'site_3', 'site_4', 'site_5', 'site_6']].values.astype('float32')
    y = df['target'].values.astype('float32')
    
    # 多进程批量处理图片
    with ProcessPoolExecutor() as executor:
        X_pic = list(executor.map(process_image, df['path'], [new_size]*len(df)))
    
    # 转换为numpy数组
    X_pic = np.array(X_pic)
    
    return (X_pic, X_stats), y

4. 预分配numpy数组(可选优化)

若明确图片最终形状,可提前分配数组,避免列表append后的数组转换开销:

# 预分配数组替代list转array
X_pic = np.zeros((len(df), new_size[1], new_size[0], 3), dtype='float32')
with ProcessPoolExecutor() as executor:
    for idx, img in enumerate(executor.map(process_image, df['path'], [new_size]*len(df))):
        X_pic[idx] = img

效果预期

以上优化后,3万行数据的处理时间可压缩至十几分钟甚至更短,具体耗时取决于CPU核心数和磁盘IO性能。

内容的提问来源于stack exchange,提问作者bilbo_slagins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 19:02:33