You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用concurrent.futures实现单文件读取后多任务并发处理?

优化并发读取文件的脚本效率问题

我想在脚本里用concurrent.futures实现异步I/O读取,目标是仅读取一次文件后再处理结果。但现在的实现是两个函数各自读取文件转成pandas DataFrame再返回,重复读了两次。之前试过在全局作用域先读文件,再在函数里用,但并发处理时出现了“Futures对象无to_dict、values[0]属性”的错误。请问怎么正确用并发/线程模块来优化脚本效率?

原始代码

import pandas as pd
import sys,os, time,re
import concurrent.futures

start=(time.perf_counter())
def getting_file_path(fileName):
    if getattr(sys, 'frozen', False) and hasattr(sys, '_MEIPASS'):
        path_actual = os.getcwd()
        path_main_folder = path_actual[:-4]
        path_result = path_main_folder + fileName
        print('frozen path',os.path.normpath(path_result))
        return path_result
    else:
        return fileName


def read_keys_dropdown():
    global lst_dropdown_keys
    file_to_read = pd.read_json(getting_file_path('./ConfigurationFile/configFile.csv'))
    lst_dropdown_keys=list(file_to_read.to_dict().keys())
    lst_dropdown_keys.pop(0)
    lst_dropdown_keys.pop(-1)
    return lst_dropdown_keys


def read_url():
    pattern = re.compile(r"^(?:/.|[^//])*/((?:\\.|[^/\\])*)/")
    file_to_read=pd.read_json(getting_file_path('./ConfigurationFile/configFile.csv'))
    result = (re.match(pattern, file_to_read.values[0][0]))
    return pattern.match(file_to_read.values[0][0]).group(1)


with concurrent.futures.ThreadPoolExecutor() as executor:
    res_1=executor.submit(read_keys_dropdown)
    res_2=executor.submit(read_url)
    finish=(time.perf_counter())
    print(res_1.result(),res_2.result(),finish-start,sep=';')

解决方案

核心优化逻辑

  • 主线程一次性完成文件读取,避免重复磁盘I/O开销
  • 将加载好的DataFrame作为参数传给并发处理函数,实现数据共享
  • 修正Future对象的使用方式:必须通过.result()获取实际处理结果后再操作

修改后的代码

import pandas as pd
import sys, os, time, re
import concurrent.futures

start = time.perf_counter()

def getting_file_path(fileName):
    if getattr(sys, 'frozen', False) and hasattr(sys, '_MEIPASS'):
        path_actual = os.getcwd()
        path_main_folder = path_actual[:-4]
        path_result = path_main_folder + fileName
        print('frozen path', os.path.normpath(path_result))
        return path_result
    else:
        return fileName

# 仅负责处理下拉菜单keys的函数,接收已加载的DataFrame
def process_dropdown_keys(df):
    lst_dropdown_keys = list(df.to_dict().keys())
    lst_dropdown_keys.pop(0)
    lst_dropdown_keys.pop(-1)
    return lst_dropdown_keys

# 仅负责处理URL的函数,接收已加载的DataFrame
def process_url(df):
    pattern = re.compile(r"^(?:/.|[^//])*/((?:\\.|[^/\\])*)/")
    match_result = pattern.match(df.values[0][0])
    return match_result.group(1)

# 主线程先读取一次文件(关键:只做一次I/O)
file_path = getting_file_path('./ConfigurationFile/configFile.csv')
df = pd.read_json(file_path)

# 线程池并发处理内存中的数据
with concurrent.futures.ThreadPoolExecutor() as executor:
    # 提交任务时传入已加载的DataFrame
    future_keys = executor.submit(process_dropdown_keys, df)
    future_url = executor.submit(process_url, df)
    
    # 获取任务执行结果(必须先调用.result()拿到实际数据)
    res_1 = future_keys.result()
    res_2 = future_url.result()

finish = time.perf_counter()
print(res_1, res_2, finish - start, sep=';')

关键说明

  1. 重复I/O消除:文件读取是磁盘操作,耗时远高于内存计算,主线程一次性加载后所有线程共享数据,直接减少一半的I/O开销
  2. 线程安全保障:pandas DataFrame在只读场景下是线程安全的,多个线程同时读取数据不会出现异常
  3. 错误根源修正:之前的“Futures对象无xxx属性”错误,是因为直接对executor.submit()返回的Future对象调用数据方法,必须先通过.result()获取实际返回值再操作

内容的提问来源于stack exchange,提问作者xlmaster

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 10:12:45