Python多重型CSV处理:线程池/进程池选型及max_workers设置
问题
我有数量不确定的CSV文件,需要先备份再执行重型处理(含行列值操作、搜索、值添加等),最后保存结果。因任务较重,希望尽可能实现并发处理。
现有两段代码分别使用ThreadPoolExecutor和ProcessPoolExecutor,不确定哪种方案更合适——任务包含文件读写,但核心函数my_heavy_process为CPU密集型,无法判断整体任务类型。同时疑惑max_workers如何设置,文件数为10或40时是否需调整?
现有代码
方案1:ThreadPoolExecutor实现
import concurrent.futures import pandas as pd import os def process_csv(file_path): # Modify the file name for backup backup_file_path = file_path.replace('.csv', '_BACK.csv') # Backup the file os.rename(file_path, backup_file_path) # Process the CSV file df = pd.read_csv(backup_file_path) my_heavy_process(df) # Save the processed DataFrame to a new file result_file_path = file_path.replace('.csv', '_result.csv') df.to_csv(result_file_path, index=False) # List of CSV files to process. It is supposed that I am going # to read it dynamically from a folder csv_files = ['file1.csv', 'file2.csv', 'file3.csv', 'file4.csv', 'file5.csv'] # Create a ThreadPoolExecutor with max_workers set to the number of files for parallel processing with concurrent.futures.ThreadPoolExecutor(max_workers=len(csv_files)) as executor: # Submit each CSV file for processing futures = {executor.submit(process_csv, file): file for file in csv_files} # Wait for all tasks to complete concurrent.futures.wait(futures)
方案2:ProcessPoolExecutor实现
import concurrent.futures import pandas as pd import os def process_csv(file_path): # Modify the file name for backup backup_file_path = file_path.replace('.csv', '_BACK.csv') # Backup the file os.rename(file_path, backup_file_path) # Process the CSV file df = pd.read_csv(backup_file_path) my_heavy_process(df) # Save the processed DataFrame to a new file result_file_path = file_path.replace('.csv', '_result.csv') df.to_csv(result_file_path, index=False) # List of CSV files to process. It is supposed that I am going # to read it dynamically from a folder csv_files = ['file1.csv', 'file2.csv', 'file3.csv', 'file4.csv', 'file5.csv'] with concurrent.futures.ProcessPoolExecutor(max_workers=len(csv_files)) as executor: # Submit each CSV file for processing futures = {executor.submit(process_csv, file): file for file in csv_files} # Wait for all tasks to complete concurrent.futures.wait(futures)
分析与解答
一、选择ThreadPoolExecutor还是ProcessPoolExecutor?
优先选ProcessPoolExecutor,原因如下:
- 你的任务核心耗时部分是CPU密集型的
my_heavy_process,而Python的GIL(全局解释器锁)会限制多线程的CPU并行效率——同一时间只有一个线程能执行Python字节码,多线程处理CPU密集任务本质是串行,无法利用多核资源。 ProcessPoolExecutor通过创建独立进程绕过GIL,每个进程有自己的Python解释器和内存空间,能真正并行利用多核CPU处理核心任务,大幅提升执行速度。- 任务中的IO密集部分(文件读写、CSV读写)开销远低于CPU处理,进程间切换的少量额外成本完全可以忽略。
二、max_workers的设置建议
不要直接设为文件总数,需结合CPU核心数调整:
- CPU密集型为主场景:
max_workers建议设为CPU核心数的1~2倍。比如8核CPU,设为8或16即可。过多进程会导致CPU上下文切换频繁,反而降低效率。 - 文件数远大于CPU核心数时(比如40个文件、8核CPU):
- 设为CPU核心数的1~2倍足够,多余任务会在队列中等待,系统自动调度,不会因进程过多拖慢整体速度。
- 强行设为40会创建大量进程,远超CPU处理能力,反而因进程切换开销增加整体耗时。
- 动态设置推荐:用
os.cpu_count()获取当前系统核心数,结合文件数做限制:
既保证不创建过多进程,也能充分利用CPU资源。max_workers = min(os.cpu_count() * 2, len(csv_files))
额外优化点
- 备份文件时,
os.rename跨磁盘分区可能失败,建议用shutil.copy2代替,保留文件元信息且支持跨分区:import shutil shutil.copy2(file_path, backup_file_path) - 给
concurrent.futures.wait添加超时参数,避免单个任务卡死导致无限等待:concurrent.futures.wait(futures, timeout=3600) # 超时1小时,可按需调整
内容的提问来源于stack exchange,提问作者KansaiRobot
相关产品推荐
相关产品推荐

