如何按文件名日期合并H5数据NumPy数组,按月+ID分组求平均
问题描述
我有数百个文件名包含日期(例如...20221017...)的.h5文件,已从每个文件提取参数并整理为如下格式的NumPy数组:
[[param_1a, param_2a...param_5a], ... [param_1x, param_2x,...param_5x]]
需要同时按ID(每行第一个数字)和月份分组,将同一月份、同一ID的参数计算平均值,合并成代表该月份的结果数组。
现有代码片段如下(filename是存储文件名的txt文件路径):
def combine_months(filename): fin = open(filename, 'r') next_name = fin.readline() while (next_name != ""): year = next_name[6:10] month = next_name[11:13] date = month+'\\'+year #not sure where to go from here fin.close()
示例说明
假设同一月份的三个文件提取出以下数组:
array_1 = [[ 1, 4, 10], [ 2, 5, 11], [3, 6, 12]] array_2 = [[ 1, 2, 5], [ 2, 2, 3], [ 3, 6, 12]] array_3 = [[ 2, 4, 10], [ 3, 2, 3], [ 4, 6, 12]]
期望得到的结果(以2022年4月为例):
2022_04_data = [[1, 3, 7.5], [2, 2, 6.5], [3, 4, 7.5], [4, 6, 12]]
解决方案
核心思路
- 用嵌套字典分层存储数据:外层键为
年_月格式(如2022_04),内层键为ID,值存储对应参数的累加和与出现次数。 - 遍历所有文件名,提取年月信息后读取对应
.h5文件的数组。 - 对数组每行按ID和年月分组,累加参数值并计数。
- 最后对每个分组计算平均值,整理为结构化的结果数组。
完整实现代码
import numpy as np from collections import defaultdict def combine_months(filename): # 嵌套字典:year_month -> id -> [参数累加和, 出现次数] # 假设每行有5个参数,可根据实际情况调整np.zeros的维度 data_groups = defaultdict(lambda: defaultdict(lambda: [np.zeros(5), 0])) with open(filename, 'r') as fin: for line in fin: file_path = line.strip() if not file_path: continue # 从文件名提取年月(注意:需根据你的文件名实际格式调整切片位置) year = file_path[6:10] month = file_path[11:13] year_month_key = f"{year}_{month}" # 读取.h5文件中的数组(请根据你的h5文件结构实现read_h5_data函数) arr = read_h5_data(file_path) # 遍历数组每行,更新分组数据 for row in arr: id_val = row[0] params = row[1:] # 提取ID以外的参数 data_groups[year_month_key][id_val][0] += params data_groups[year_month_key][id_val][1] += 1 # 生成最终的月度平均数组 monthly_results = {} for year_month, id_records in data_groups.items(): avg_rows = [] for id_val, (sum_params, count) in id_records.items(): avg_params = sum_params / count avg_rows.append([id_val] + avg_params.tolist()) # 按ID排序结果(可选,按需移除) avg_rows.sort(key=lambda x: x[0]) monthly_results[year_month] = np.array(avg_rows) return monthly_results # 示例:读取.h5文件数据的函数(需根据你的h5文件实际存储结构修改) def read_h5_data(file_path): import h5py with h5py.File(file_path, 'r') as hf: # 假设数据存储在'dataset'键下,替换为你实际的键名 return hf['dataset'][:] # 使用示例 if __name__ == "__main__": results = combine_months("filenames.txt") # 输出2022年4月的结果 if "2022_04" in results: print(results["2022_04"])
代码说明
- 内存优化:通过累加参数和计数的方式,避免存储所有行数据,适合处理数百个文件的大规模数据。
- 灵活性:参数数量可通过修改
np.zeros(5)调整,文件名的年月提取逻辑可根据实际格式修改切片位置。 - 结构化输出:结果以字典形式返回,键为
年_月,值为对应月份的平均数组,方便后续调用和分析。
内容的提问来源于stack exchange,提问作者brizzy
相关产品推荐
相关产品推荐

