如何使用concurrent.futures实现单文件读取后多任务并发处理?
优化并发读取文件的脚本效率问题
我想在脚本里用concurrent.futures实现异步I/O读取,目标是仅读取一次文件后再处理结果。但现在的实现是两个函数各自读取文件转成pandas DataFrame再返回,重复读了两次。之前试过在全局作用域先读文件,再在函数里用,但并发处理时出现了“Futures对象无to_dict、values[0]属性”的错误。请问怎么正确用并发/线程模块来优化脚本效率?
原始代码
import pandas as pd import sys,os, time,re import concurrent.futures start=(time.perf_counter()) def getting_file_path(fileName): if getattr(sys, 'frozen', False) and hasattr(sys, '_MEIPASS'): path_actual = os.getcwd() path_main_folder = path_actual[:-4] path_result = path_main_folder + fileName print('frozen path',os.path.normpath(path_result)) return path_result else: return fileName def read_keys_dropdown(): global lst_dropdown_keys file_to_read = pd.read_json(getting_file_path('./ConfigurationFile/configFile.csv')) lst_dropdown_keys=list(file_to_read.to_dict().keys()) lst_dropdown_keys.pop(0) lst_dropdown_keys.pop(-1) return lst_dropdown_keys def read_url(): pattern = re.compile(r"^(?:/.|[^//])*/((?:\\.|[^/\\])*)/") file_to_read=pd.read_json(getting_file_path('./ConfigurationFile/configFile.csv')) result = (re.match(pattern, file_to_read.values[0][0])) return pattern.match(file_to_read.values[0][0]).group(1) with concurrent.futures.ThreadPoolExecutor() as executor: res_1=executor.submit(read_keys_dropdown) res_2=executor.submit(read_url) finish=(time.perf_counter()) print(res_1.result(),res_2.result(),finish-start,sep=';')
解决方案
核心优化逻辑
- 主线程一次性完成文件读取,避免重复磁盘I/O开销
- 将加载好的DataFrame作为参数传给并发处理函数,实现数据共享
- 修正Future对象的使用方式:必须通过
.result()获取实际处理结果后再操作
修改后的代码
import pandas as pd import sys, os, time, re import concurrent.futures start = time.perf_counter() def getting_file_path(fileName): if getattr(sys, 'frozen', False) and hasattr(sys, '_MEIPASS'): path_actual = os.getcwd() path_main_folder = path_actual[:-4] path_result = path_main_folder + fileName print('frozen path', os.path.normpath(path_result)) return path_result else: return fileName # 仅负责处理下拉菜单keys的函数,接收已加载的DataFrame def process_dropdown_keys(df): lst_dropdown_keys = list(df.to_dict().keys()) lst_dropdown_keys.pop(0) lst_dropdown_keys.pop(-1) return lst_dropdown_keys # 仅负责处理URL的函数,接收已加载的DataFrame def process_url(df): pattern = re.compile(r"^(?:/.|[^//])*/((?:\\.|[^/\\])*)/") match_result = pattern.match(df.values[0][0]) return match_result.group(1) # 主线程先读取一次文件(关键:只做一次I/O) file_path = getting_file_path('./ConfigurationFile/configFile.csv') df = pd.read_json(file_path) # 线程池并发处理内存中的数据 with concurrent.futures.ThreadPoolExecutor() as executor: # 提交任务时传入已加载的DataFrame future_keys = executor.submit(process_dropdown_keys, df) future_url = executor.submit(process_url, df) # 获取任务执行结果(必须先调用.result()拿到实际数据) res_1 = future_keys.result() res_2 = future_url.result() finish = time.perf_counter() print(res_1, res_2, finish - start, sep=';')
关键说明
- 重复I/O消除:文件读取是磁盘操作,耗时远高于内存计算,主线程一次性加载后所有线程共享数据,直接减少一半的I/O开销
- 线程安全保障:pandas DataFrame在只读场景下是线程安全的,多个线程同时读取数据不会出现异常
- 错误根源修正:之前的“Futures对象无xxx属性”错误,是因为直接对
executor.submit()返回的Future对象调用数据方法,必须先通过.result()获取实际返回值再操作
内容的提问来源于stack exchange,提问作者xlmaster
相关产品推荐
相关产品推荐

