You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python高效删除LAMMPS模拟dump文件中的特定行?

LAMMPS Dump文件原子数据提取优化建议

我用LAMMPS进行模拟得到的dump文件中,每个时间步前有9行仅含信息的非数据行,需要删除这些行以提取纯原子数据并保存至独立文件。目前已有实现该功能的Python代码,但作为Python新手,希望获取优化建议提升处理速度。

Dump文件片段

ITEM: TIMESTEP
0
ITEM: NUMBER OF ATOMS
4200
ITEM: BOX BOUNDS pp pp pp
-2.0000000000000000e+01 2.0000000000000000e+01
-2.0000000000000000e+01 2.0000000000000000e+01
-2.0000000000000000e+01 2.0000000000000000e+01
ITEM: ATOMS id mol xu yu zu
533 26 -17.891 -16.7503 -18.8102
534 26 -17.7164 -17.5276 -18.7004
535 26 -17.3612 -17.7508 -19.2693
536 26 -17.0213 -17.8009 -18.5118
537 26 -17.8409 -18.5307 -18.8511
538 26 -17.7968 -19.5713 -18.6246
ITEM: TIMESTEP
1
ITEM: NUMBER OF ATOMS
4200
ITEM: BOX BOUNDS pp pp pp
-2.0000000000000000e+01 2.0000000000000000e+01
-2.0000000000000000e+01 2.0000000000000000e+01
-2.0000000000000000e+01 2.0000000000000000e+01
ITEM: ATOMS id mol xu yu zu
536 26 -17.0213 -17.8009 -18.5118
537 26 -17.8409 -18.5307 -18.8511
538 26 -17.7968 -19.5713 -18.6246

当前代码

def data_process_func(filename, n_atoms, k):
    
    with open(filename, 'r') as f:
        lines = f.readlines()
    
    # The following loop deletes all the text only leaving data
    for i in range(len(timestep)):
        del lines[(n_atoms)*i:(n_atoms*i)+9]
        
    # Saves the data without the text to a txt file
    with open('data_{}.txt'.format(k), 'w') as f:
        f.writelines(lines)

    # Loads the data from the file into a dataframe
    data = pd.read_csv('data_{}.txt'.format(k), sep=" ", header = None, names = ['id', 'mol', 'xu', 'yu', 'zu'])
        
    return data

优化建议

1. 先修复原代码的致命问题

原代码存在两个严重bug:

  • len(timestep)中的timestep未定义,运行会直接报错;
  • 循环删除列表元素时,列表长度会动态变化,后续的索引计算完全错位,无法正确删除目标行。

2. 改用逐行过滤写入,减少内存占用

原代码先全量读入所有行再删除,大文件会占用大量内存且效率低。改为逐行读取,判断当前行是否属于原子数据行,直接写入目标文件,无需存储所有行:

import pandas as pd

def data_process_func(filename, n_atoms, k):
    output_file = f'data_{k}.txt'
    # 同时打开输入和输出文件,逐行处理
    with open(filename, 'r') as f_in, open(output_file, 'w') as f_out:
        line_idx = 0
        # 每个时间块包含9行头 + n_atoms行数据
        block_size = 9 + n_atoms
        for line in f_in:
            # 跳过每个块的前9行头信息
            if line_idx % block_size >= 9:
                f_out.write(line)
            line_idx += 1
    # 读取生成的文件到DataFrame
    data = pd.read_csv(output_file, sep=" ", header=None, names=['id', 'mol', 'xu', 'yu', 'zu'])
    return data

3. 跳过中间文件,直接生成DataFrame

省去写入再读取txt文件的IO开销,直接将符合条件的行解析后存入列表,再生成DataFrame:

import pandas as pd

def data_process_func(filename, n_atoms, k):
    data_rows = []
    with open(filename, 'r') as f_in:
        line_idx = 0
        block_size = 9 + n_atoms
        for line in f_in:
            if line_idx % block_size >= 9:
                # 拆分每行数据并转换对应类型
                parts = line.strip().split()
                data_rows.append([
                    int(parts[0]), int(parts[1]),
                    float(parts[2]), float(parts[3]), float(parts[4])
                ])
            line_idx += 1
    # 直接生成DataFrame
    data = pd.DataFrame(data_rows, columns=['id', 'mol', 'xu', 'yu', 'zu'])
    # 如果需要保存文件,再执行写入
    data.to_csv(f'data_{k}.txt', sep=' ', header=False, index=False)
    return data

4. 超大文件场景:用生成器降低内存占用

如果dump文件特别大(几十GB级别),用生成器逐行解析数据,避免将所有数据存入内存:

import pandas as pd

def atom_data_generator(filename, n_atoms):
    with open(filename, 'r') as f_in:
        line_idx = 0
        block_size = 9 + n_atoms
        for line in f_in:
            if line_idx % block_size >= 9:
                parts = line.strip().split()
                yield (
                    int(parts[0]), int(parts[1]),
                    float(parts[2]), float(parts[3]), float(parts[4])
                )
            line_idx += 1

def data_process_func(filename, n_atoms, k):
    # 用生成器逐行提供数据
    data_gen = atom_data_generator(filename, n_atoms)
    data = pd.DataFrame(data_gen, columns=['id', 'mol', 'xu', 'yu', 'zu'])
    data.to_csv(f'data_{k}.txt', sep=' ', header=False, index=False)
    return data

内容的提问来源于stack exchange,提问作者Morten jørgensen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 13:10:26