读取文本文件校正列间距、处理重复头并写入HDF5的问题咨询
GROMACS轨迹文件转HDF5格式问题解决方案
报错原因
- 当前代码采用空格分隔读取gro格式文件,但gro为固定宽度格式,原子序号超过4位后第2、3列的分隔空格会消失,导致识别为5列触发报错
- 后续重复出现的帧头行也会被识别为非法数据行,进一步触发列数不匹配错误
解决方案
1. 采用固定宽度切片读取数据
直接按gro文件的固定列宽切割每行内容,不需要依赖空格分隔,从根源解决列拼接问题:
# 定义每行各字段的切片规则 def parse_line(line): col1 = line[0:10].strip() col2_raw = line[10:20].strip() # 拆分拼接的col2和col3:col2是前4个字符,剩余部分是col3的数字 col2 = col2_raw[:4] col3 = int(col2_raw[4:]) col4 = float(line[20:28].strip()) col5 = float(line[28:36].strip()) col6 = float(line[36:44].strip()) return (col1.encode('utf-8'), col2.encode('utf-8'), col3, col4, col5, col6)
2. 自动识别帧头处理多帧数据
每帧结构固定为「1行帧头 + 1行原子数 + N行原子数据」,逐行扫描识别帧头即可,不需要提前知道帧头行号,同时避免一次性读入全量文件导致内存溢出:
import h5py import numpy as np gro_dt = np.dtype([('col1', 'S4'), ('col2', 'S4'), ('col3', int), ('col4', float), ('col5', float), ('col6', float)]) with h5py.File('xaa.h5', 'w') as hdf, open('xaa', 'r') as f: # 主数据集存所有原子数据 ds = hdf.create_dataset('dataset1', dtype=gro_dt, shape=(0,), maxshape=(None,)) # 单独存每个帧的时间戳和帧起始索引,方便后续按帧检索 time_ds = hdf.create_dataset('timesteps', dtype=np.float64, shape=(0,), maxshape=(None,)) frame_idx_ds = hdf.create_dataset('frame_start_idx', dtype=np.int64, shape=(0,), maxshape=(None,)) # 若需要存储完整帧头字符串,可新增如下字符串数据集 str_dt = h5py.string_dtype(encoding='utf-8') header_ds = hdf.create_dataset('frame_headers', dtype=str_dt, shape=(0,), maxshape=(None,)) row_count = 0 while True: line = f.readline() if not line: break # 识别帧头 if line.startswith('Generated by trjconv'): # 写入完整帧头 header_ds.resize((header_ds.shape[0]+1,)) header_ds[-1] = line.strip() # 提取时间戳 t = float(line.strip().split('t=')[-1]) # 读下一行的原子数 n_atom = int(f.readline().strip()) # 记录当前帧的索引和时间 time_ds.resize((time_ds.shape[0]+1,)) time_ds[-1] = t frame_idx_ds.resize((frame_idx_ds.shape[0]+1,)) frame_idx_ds[-1] = row_count # 读取当前帧的所有原子数据 frame_data = [] for _ in range(n_atom): atom_line = f.readline() frame_data.append(parse_line(atom_line)) # 写入主数据集 ds.resize((row_count + n_atom,)) ds[row_count:row_count+n_atom] = np.array(frame_data, dtype=gro_dt) row_count += n_atom
3. HDF5存储字符串的说明
HDF5完全支持字符串存储,除了上述单独创建字符串数据集的方式外,也可以将帧头信息存储为数据集的属性,按需选择即可。如果不需要保留原始帧头全文,仅提取时间戳存储为浮点数是性能最优、占用空间最小的方案。
内容的提问来源于stack exchange,提问作者Mahesh
相关产品推荐
相关产品推荐

