You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从数据文件按标识筛选行并拆分存入HDF5不同分组

实现方法

你不需要拆分原始文件,只需要在逐行解析数据时加标识判断,分别存储两类数据即可,修改逻辑如下:

  • 逐行解析时提取行首ID字段,判断是含P的蛋白数据还是含LI的脂质数据,分别存入两个临时数组
  • 初始化HDF5数据集时,分别为蛋白、脂质数据集配置正确的单帧长度(14、11200)
  • 每帧解析完成后,分别将两类数据写入对应的数据集
修改后的完整代码
#!/usr/bin/env python
# -*- coding: utf-8 -*-

import struct
import numpy as np
import h5py
import re

csv_file = 'com'
fmtstring = '7s 8s 5s 7s 7s 7s'
fieldstruct = struct.Struct(fmtstring)
parse = fieldstruct.unpack_from
# Format for footer
fmtstring1 = '10s 10s 10s'
fieldstruct1 = struct.Struct(fmtstring1)
parse1 = fieldstruct1.unpack_from

with open(csv_file, 'r') as f, \
    h5py.File('test.h5', 'w') as hdf:
    ## Particles group with the attributes
    particles_grp = hdf.require_group('particles/lipids/box')
    particles_grp.attrs['dimension'] = 3
    particles_grp.attrs['boundary'] = ['periodic', 'periodic', 'periodic']
    pos_grp = particles_grp.require_group('positions')
    edge_grp = particles_grp.require_group('edges')
    ## h5md group with the attributes
    h5md_grp = hdf.require_group('h5md')
    h5md_grp.attrs['version'] = 1.0
    author_grp = h5md_grp.require_group('author')
    author_grp.attrs['author'] = 'foo', 'email=foo@googlemail.com'
    creator_grp = h5md_grp.require_group('creator')
    creator_grp.attrs['name'] = 'foo'
    creator_grp.attrs['version'] = 1.0
    # datasets with known sizes
    ds_time = pos_grp.create_dataset('time', dtype="f", shape=(0,),
                                           maxshape=(None,), compression='gzip', 
                                           shuffle=True)
    ds_step = pos_grp.create_dataset('step', dtype=np.uint64, shape=(0,),
                                           maxshape=(None,), compression='gzip',
                                           shuffle=True)
    ds_protein = None
    ds_lipid = None
    # datasets in edge group
    edge_ds_time = edge_grp.create_dataset('time', dtype="f", shape=(0,),
                                           maxshape=(None,), compression='gzip', 
                                           shuffle=True)
    edge_ds_step = edge_grp.create_dataset('step', dtype="f", shape=(0,),
                                           maxshape=(None,), compression='gzip', 
                                           shuffle=True)
    edge_ds_value = None

    edge_data = edge_grp.require_dataset('box_size', dtype=np.float32, shape=(0,3),
                                              maxshape=(None,3), compression='gzip', 
                                              shuffle=True)
    
    step = 0
    while True:
        header = f.readline()
        m = re.search("t= *(.*)$", header)
        if m:
            time = float(m.group(1))
        else:
            print("End Of File")
            break

        # get number of data rows, i.e., number of particles
        nparticles = int(f.readline())
        # 初始化两个临时数组分别存蛋白和脂质数据
        protein_arr = np.empty(shape=(14, 3), dtype=np.float32)
        lipid_arr = np.empty(shape=(11200, 3), dtype=np.float32)
        p_idx = 0
        l_idx = 0
        for row in range(nparticles):
            line_bytes = f.readline().encode('utf-8')
            fields = parse(line_bytes)
            # 提取行首ID判断类型
            particle_id = fields[0].decode('utf-8').strip()
            pos = np.array((float(fields[3]), float(fields[4]), float(fields[5])))
            if 'P' in particle_id:
                protein_arr[p_idx] = pos
                p_idx += 1
            elif 'LI' in particle_id:
                lipid_arr[l_idx] = pos
                l_idx += 1
        if nparticles > 0:
            # 首次迭代创建两个数据集
            if ds_protein is None:
                # 蛋白数据集:单帧14条
                ds_protein = pos_grp.create_dataset('protein', dtype=np.float32,
                                                        shape=(0, 14, 3), maxshape=(None, 14, 3),
                                                        chunks=(1, 14, 3), compression='gzip', shuffle=True)
                # 脂质数据集:单帧11200条
                ds_lipid = pos_grp.create_dataset('lipid', dtype=np.float32,
                                                        shape=(0, 11200, 3), maxshape=(None, 11200, 3),
                                                        chunks=(1, 11200, 3), compression='gzip', shuffle=True)
                edge_ds_value = edge_grp.create_dataset('value', dtype=np.float32,
                                                        shape=(0, 3), maxshape=(None, 3),chunks=(1, 3), compression='gzip', shuffle=True)
            # 公共维度扩容
            ds_time.resize(step + 1, axis=0)
            ds_step.resize(step + 1, axis=0)
            # 两个数据集分别扩容
            ds_protein.resize(step + 1, axis=0)
            ds_lipid.resize(step + 1, axis=0)
            # edge组扩容
            edge_ds_time.resize(step + 1, axis=0)
            edge_ds_step.resize(step + 1, axis=0)
            edge_ds_value.resize(step + 1, axis=0)
            
            # 赋值写入
            ds_time[step] = time
            ds_step[step] = step
            ds_protein[step] = protein_arr
            ds_lipid[step] = lipid_arr
            edge_ds_time[step] = time
            edge_ds_step[step] = step

        footer = parse1( f.readline().encode('utf-8') )
        dat = np.array(footer).astype(float).reshape(1,3)
        new_size = edge_data.shape[0] + 1
        edge_data.resize(new_size, axis=0)
        edge_data[new_size-1 : new_size, :] = dat
        step += 1
关键修改说明
  • 原代码中只创建了一个数据集存储全量数据,修改后分别创建protein和lipid两个数据集,对应两类数据的单帧长度
  • 逐行解析时增加了ID判断逻辑,将不同类型的坐标存入对应临时数组
  • 修复了原代码中蛋白数据集被错误命名为lipid的问题

内容的提问来源于stack exchange,提问作者Mahesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 21:54:03