如何从数据文件按标识筛选行并拆分存入HDF5不同分组
实现方法
你不需要拆分原始文件,只需要在逐行解析数据时加标识判断,分别存储两类数据即可,修改逻辑如下:
- 逐行解析时提取行首ID字段,判断是含
P的蛋白数据还是含LI的脂质数据,分别存入两个临时数组 - 初始化HDF5数据集时,分别为蛋白、脂质数据集配置正确的单帧长度(14、11200)
- 每帧解析完成后,分别将两类数据写入对应的数据集
修改后的完整代码
#!/usr/bin/env python # -*- coding: utf-8 -*- import struct import numpy as np import h5py import re csv_file = 'com' fmtstring = '7s 8s 5s 7s 7s 7s' fieldstruct = struct.Struct(fmtstring) parse = fieldstruct.unpack_from # Format for footer fmtstring1 = '10s 10s 10s' fieldstruct1 = struct.Struct(fmtstring1) parse1 = fieldstruct1.unpack_from with open(csv_file, 'r') as f, \ h5py.File('test.h5', 'w') as hdf: ## Particles group with the attributes particles_grp = hdf.require_group('particles/lipids/box') particles_grp.attrs['dimension'] = 3 particles_grp.attrs['boundary'] = ['periodic', 'periodic', 'periodic'] pos_grp = particles_grp.require_group('positions') edge_grp = particles_grp.require_group('edges') ## h5md group with the attributes h5md_grp = hdf.require_group('h5md') h5md_grp.attrs['version'] = 1.0 author_grp = h5md_grp.require_group('author') author_grp.attrs['author'] = 'foo', 'email=foo@googlemail.com' creator_grp = h5md_grp.require_group('creator') creator_grp.attrs['name'] = 'foo' creator_grp.attrs['version'] = 1.0 # datasets with known sizes ds_time = pos_grp.create_dataset('time', dtype="f", shape=(0,), maxshape=(None,), compression='gzip', shuffle=True) ds_step = pos_grp.create_dataset('step', dtype=np.uint64, shape=(0,), maxshape=(None,), compression='gzip', shuffle=True) ds_protein = None ds_lipid = None # datasets in edge group edge_ds_time = edge_grp.create_dataset('time', dtype="f", shape=(0,), maxshape=(None,), compression='gzip', shuffle=True) edge_ds_step = edge_grp.create_dataset('step', dtype="f", shape=(0,), maxshape=(None,), compression='gzip', shuffle=True) edge_ds_value = None edge_data = edge_grp.require_dataset('box_size', dtype=np.float32, shape=(0,3), maxshape=(None,3), compression='gzip', shuffle=True) step = 0 while True: header = f.readline() m = re.search("t= *(.*)$", header) if m: time = float(m.group(1)) else: print("End Of File") break # get number of data rows, i.e., number of particles nparticles = int(f.readline()) # 初始化两个临时数组分别存蛋白和脂质数据 protein_arr = np.empty(shape=(14, 3), dtype=np.float32) lipid_arr = np.empty(shape=(11200, 3), dtype=np.float32) p_idx = 0 l_idx = 0 for row in range(nparticles): line_bytes = f.readline().encode('utf-8') fields = parse(line_bytes) # 提取行首ID判断类型 particle_id = fields[0].decode('utf-8').strip() pos = np.array((float(fields[3]), float(fields[4]), float(fields[5]))) if 'P' in particle_id: protein_arr[p_idx] = pos p_idx += 1 elif 'LI' in particle_id: lipid_arr[l_idx] = pos l_idx += 1 if nparticles > 0: # 首次迭代创建两个数据集 if ds_protein is None: # 蛋白数据集:单帧14条 ds_protein = pos_grp.create_dataset('protein', dtype=np.float32, shape=(0, 14, 3), maxshape=(None, 14, 3), chunks=(1, 14, 3), compression='gzip', shuffle=True) # 脂质数据集:单帧11200条 ds_lipid = pos_grp.create_dataset('lipid', dtype=np.float32, shape=(0, 11200, 3), maxshape=(None, 11200, 3), chunks=(1, 11200, 3), compression='gzip', shuffle=True) edge_ds_value = edge_grp.create_dataset('value', dtype=np.float32, shape=(0, 3), maxshape=(None, 3),chunks=(1, 3), compression='gzip', shuffle=True) # 公共维度扩容 ds_time.resize(step + 1, axis=0) ds_step.resize(step + 1, axis=0) # 两个数据集分别扩容 ds_protein.resize(step + 1, axis=0) ds_lipid.resize(step + 1, axis=0) # edge组扩容 edge_ds_time.resize(step + 1, axis=0) edge_ds_step.resize(step + 1, axis=0) edge_ds_value.resize(step + 1, axis=0) # 赋值写入 ds_time[step] = time ds_step[step] = step ds_protein[step] = protein_arr ds_lipid[step] = lipid_arr edge_ds_time[step] = time edge_ds_step[step] = step footer = parse1( f.readline().encode('utf-8') ) dat = np.array(footer).astype(float).reshape(1,3) new_size = edge_data.shape[0] + 1 edge_data.resize(new_size, axis=0) edge_data[new_size-1 : new_size, :] = dat step += 1
关键修改说明
- 原代码中只创建了一个数据集存储全量数据,修改后分别创建
protein和lipid两个数据集,对应两类数据的单帧长度 - 逐行解析时增加了ID判断逻辑,将不同类型的坐标存入对应临时数组
- 修复了原代码中蛋白数据集被错误命名为
lipid的问题
内容的提问来源于stack exchange,提问作者Mahesh
相关产品推荐
相关产品推荐

