优化Python函数运行效率:解决重复读取同一文件的问题
优化方案:避免重复读取同一文件
核心优化思路:先按文件路径分组关联的变量名,让每个文件仅被读取和计算一次,再把结果批量赋值给所有对应的变量列,彻底消除重复IO和计算开销。
步骤1:反转原字典,构建文件与变量的映射
使用collections.defaultdict将原“变量→文件”的字典反转,变成“文件→关联变量列表”的结构,方便按文件批量处理:
from collections import defaultdict FILEPATH = {"variable_1": "path/commonfile.tif", "variable_2": "path/commonfile.tif", "variable_3": "path/commonfile.tif", "variable_4": "path/otherfile1.tif", "variable_5": "path/someotherfile1.tif"} # 反转字典,按文件路径分组变量 file_to_vars = defaultdict(list) for var, path in FILEPATH.items(): file_to_vars[path].append(var)
处理后file_to_vars的结构为:
{ "path/commonfile.tif": ["variable_1", "variable_2", "variable_3"], "path/otherfile1.tif": ["variable_4"], "path/someotherfile1.tif": ["variable_5"] }
步骤2:批量处理文件,复用计算结果
遍历反转后的字典,每个文件仅调用一次myfunction读取计算,再将结果批量赋值给所有关联的变量列:
# 缓存已计算的结果,避免重复处理同一文件(复杂场景下更实用) result_cache = {} for filename, variables in file_to_vars.items(): if filename not in result_cache: # 仅当文件未处理过时,执行读取和统计计算 result_cache[filename] = myfunction(df=my_df, file=filename, stats_list=['mean']) # 将结果同步赋值给所有关联的变量列 for var in variables: my_df.loc[:, var] = result_cache[filename]
完整优化代码
from collections import defaultdict FILEPATH = {"variable_1": "path/commonfile.tif", "variable_2": "path/commonfile.tif", "variable_3": "path/commonfile.tif", "variable_4": "path/otherfile1.tif", "variable_5": "path/someotherfile1.tif"} # 构建文件到变量的映射 file_to_vars = defaultdict(list) for var, path in FILEPATH.items(): file_to_vars[path].append(var) result_cache = {} for filename, variables in file_to_vars.items(): if filename not in result_cache: result_cache[filename] = myfunction(df=my_df, file=filename, stats_list=['mean']) for var in variables: my_df.loc[:, var] = result_cache[filename]
内容的提问来源于stack exchange,提问作者A.N.
相关产品推荐
相关产品推荐

