You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何并行运行37k行数据集处理任务以提速?附代码及问题

大规模数据集处理提速需求

我正在处理一个包含37k行的大规模数据集,需对其应用自定义函数caption_from_image_file,希望通过多线程或其他方式提升处理速度。此前尝试过Spark和Dask库,但遇到了无法解决的错误,以下是我的代码:

import matplotlib.pyplot as plt

def caption_from_image_file(x):
    return [str(get_caption(i,device)) for i in x.load()]

import cv2
import numpy as np

df = dg.getData("train")
df_test = df

# 启动计时器
import time
start_time = time.time()

df_test['captions'] = df_test.images.apply(caption_from_image_file)

# 结束计时(单位:分钟)
print("--- %s minutes ---" % ((time.time() - start_time)/60))

df_test.to_csv('test.csv',index=False)

# 释放CUDA内存
torch.cuda.empty_cache()

df_test.captions

内容的提问来源于stack exchange,提问作者AymaneElmahi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 15:00:49