You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多重型CSV处理:线程池/进程池选型及max_workers设置

问题

我有数量不确定的CSV文件,需要先备份再执行重型处理(含行列值操作、搜索、值添加等),最后保存结果。因任务较重,希望尽可能实现并发处理。

现有两段代码分别使用ThreadPoolExecutor和ProcessPoolExecutor,不确定哪种方案更合适——任务包含文件读写,但核心函数my_heavy_process为CPU密集型,无法判断整体任务类型。同时疑惑max_workers如何设置,文件数为10或40时是否需调整?


现有代码

方案1:ThreadPoolExecutor实现

import concurrent.futures
import pandas as pd
import os

def process_csv(file_path):
    # Modify the file name for backup
    backup_file_path = file_path.replace('.csv', '_BACK.csv')
    
    # Backup the file
    os.rename(file_path, backup_file_path)
    
    # Process the CSV file 
    df = pd.read_csv(backup_file_path)
    my_heavy_process(df)
    
    # Save the processed DataFrame to a new file
    result_file_path = file_path.replace('.csv', '_result.csv')
    df.to_csv(result_file_path, index=False)

# List of CSV files to process. It is supposed that I am going
# to read it dynamically from a folder
csv_files = ['file1.csv', 'file2.csv', 'file3.csv', 'file4.csv', 'file5.csv']

# Create a ThreadPoolExecutor with max_workers set to the number of files for parallel processing
with concurrent.futures.ThreadPoolExecutor(max_workers=len(csv_files)) as executor:
    # Submit each CSV file for processing
    futures = {executor.submit(process_csv, file): file for file in csv_files}

    # Wait for all tasks to complete
    concurrent.futures.wait(futures)

方案2:ProcessPoolExecutor实现

import concurrent.futures
import pandas as pd
import os

def process_csv(file_path):
    # Modify the file name for backup
    backup_file_path = file_path.replace('.csv', '_BACK.csv')
    
    # Backup the file
    os.rename(file_path, backup_file_path)
    
    # Process the CSV file 
    df = pd.read_csv(backup_file_path)
    my_heavy_process(df)
    
    # Save the processed DataFrame to a new file
    result_file_path = file_path.replace('.csv', '_result.csv')
    df.to_csv(result_file_path, index=False)

# List of CSV files to process. It is supposed that I am going
# to read it dynamically from a folder
csv_files = ['file1.csv', 'file2.csv', 'file3.csv', 'file4.csv', 'file5.csv']

with concurrent.futures.ProcessPoolExecutor(max_workers=len(csv_files)) as executor:
    # Submit each CSV file for processing
    futures = {executor.submit(process_csv, file): file for file in csv_files}

    # Wait for all tasks to complete
    concurrent.futures.wait(futures)

分析与解答

一、选择ThreadPoolExecutor还是ProcessPoolExecutor?

优先选ProcessPoolExecutor,原因如下:

  • 你的任务核心耗时部分是CPU密集型的my_heavy_process,而Python的GIL(全局解释器锁)会限制多线程的CPU并行效率——同一时间只有一个线程能执行Python字节码,多线程处理CPU密集任务本质是串行,无法利用多核资源。
  • ProcessPoolExecutor通过创建独立进程绕过GIL,每个进程有自己的Python解释器和内存空间,能真正并行利用多核CPU处理核心任务,大幅提升执行速度。
  • 任务中的IO密集部分(文件读写、CSV读写)开销远低于CPU处理,进程间切换的少量额外成本完全可以忽略。

二、max_workers的设置建议

不要直接设为文件总数,需结合CPU核心数调整:

  1. CPU密集型为主场景:max_workers建议设为CPU核心数的1~2倍。比如8核CPU,设为8或16即可。过多进程会导致CPU上下文切换频繁,反而降低效率。
  2. 文件数远大于CPU核心数时(比如40个文件、8核CPU):
    • 设为CPU核心数的1~2倍足够,多余任务会在队列中等待,系统自动调度,不会因进程过多拖慢整体速度。
    • 强行设为40会创建大量进程,远超CPU处理能力,反而因进程切换开销增加整体耗时。
  3. 动态设置推荐:用os.cpu_count()获取当前系统核心数,结合文件数做限制:
    max_workers = min(os.cpu_count() * 2, len(csv_files))
    
    既保证不创建过多进程,也能充分利用CPU资源。

额外优化点

  • 备份文件时,os.rename跨磁盘分区可能失败,建议用shutil.copy2代替,保留文件元信息且支持跨分区:
    import shutil
    shutil.copy2(file_path, backup_file_path)
    
  • 给concurrent.futures.wait添加超时参数,避免单个任务卡死导致无限等待:
    concurrent.futures.wait(futures, timeout=3600)  # 超时1小时,可按需调整
    

内容的提问来源于stack exchange,提问作者KansaiRobot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 08:30:28