You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按文件名日期合并H5数据NumPy数组,按月+ID分组求平均

问题描述

我有数百个文件名包含日期(例如...20221017...)的.h5文件,已从每个文件提取参数并整理为如下格式的NumPy数组:

[[param_1a, param_2a...param_5a],
 ... 
 [param_1x, param_2x,...param_5x]] 

需要同时按ID(每行第一个数字)和月份分组,将同一月份、同一ID的参数计算平均值,合并成代表该月份的结果数组。

现有代码片段如下(filename是存储文件名的txt文件路径):

def combine_months(filename):
    fin = open(filename, 'r')
    next_name = fin.readline()
    while (next_name != ""):
        year = next_name[6:10]
        month = next_name[11:13]
        date = month+'\\'+year
        #not sure where to go from here
    fin.close()

示例说明

假设同一月份的三个文件提取出以下数组:

array_1 = [[ 1,  4, 10],
           [ 2,  5, 11],
           [3,  6, 12]]
array_2 = [[ 1,  2, 5],
           [ 2,  2, 3],
           [ 3,  6, 12]]
array_3 = [[ 2,  4, 10],
           [ 3,  2, 3],
           [ 4,  6, 12]]

期望得到的结果(以2022年4月为例):

2022_04_data = [[1, 3, 7.5],
                [2, 2, 6.5],
                [3, 4, 7.5],
                [4, 6, 12]]

解决方案

核心思路

  1. 用嵌套字典分层存储数据:外层键为年_月格式(如2022_04),内层键为ID,值存储对应参数的累加和与出现次数。
  2. 遍历所有文件名,提取年月信息后读取对应.h5文件的数组。
  3. 对数组每行按ID和年月分组,累加参数值并计数。
  4. 最后对每个分组计算平均值,整理为结构化的结果数组。

完整实现代码

import numpy as np
from collections import defaultdict

def combine_months(filename):
    # 嵌套字典:year_month -> id -> [参数累加和, 出现次数]
    # 假设每行有5个参数,可根据实际情况调整np.zeros的维度
    data_groups = defaultdict(lambda: defaultdict(lambda: [np.zeros(5), 0]))
    
    with open(filename, 'r') as fin:
        for line in fin:
            file_path = line.strip()
            if not file_path:
                continue
            
            # 从文件名提取年月(注意:需根据你的文件名实际格式调整切片位置)
            year = file_path[6:10]
            month = file_path[11:13]
            year_month_key = f"{year}_{month}"
            
            # 读取.h5文件中的数组(请根据你的h5文件结构实现read_h5_data函数)
            arr = read_h5_data(file_path)
            
            # 遍历数组每行,更新分组数据
            for row in arr:
                id_val = row[0]
                params = row[1:]  # 提取ID以外的参数
                
                data_groups[year_month_key][id_val][0] += params
                data_groups[year_month_key][id_val][1] += 1
    
    # 生成最终的月度平均数组
    monthly_results = {}
    for year_month, id_records in data_groups.items():
        avg_rows = []
        for id_val, (sum_params, count) in id_records.items():
            avg_params = sum_params / count
            avg_rows.append([id_val] + avg_params.tolist())
        # 按ID排序结果(可选,按需移除)
        avg_rows.sort(key=lambda x: x[0])
        monthly_results[year_month] = np.array(avg_rows)
    
    return monthly_results

# 示例:读取.h5文件数据的函数(需根据你的h5文件实际存储结构修改)
def read_h5_data(file_path):
    import h5py
    with h5py.File(file_path, 'r') as hf:
        # 假设数据存储在'dataset'键下,替换为你实际的键名
        return hf['dataset'][:]

# 使用示例
if __name__ == "__main__":
    results = combine_months("filenames.txt")
    # 输出2022年4月的结果
    if "2022_04" in results:
        print(results["2022_04"])

代码说明

  • 内存优化:通过累加参数和计数的方式,避免存储所有行数据,适合处理数百个文件的大规模数据。
  • 灵活性:参数数量可通过修改np.zeros(5)调整,文件名的年月提取逻辑可根据实际格式修改切片位置。
  • 结构化输出:结果以字典形式返回,键为年_月,值为对应月份的平均数组,方便后续调用和分析。

内容的提问来源于stack exchange,提问作者brizzy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 09:05:24