You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

高效将海量分子构型的NumPy坐标数组转换为XYZ文件

高效生成大规模构型的XYZ文件(双原子二聚体)

问题背景

我生成了双原子分子二聚体(包含Al1、F1、Al2、F2四个原子)的33177600组构型坐标,每个原子的x、y、z坐标分别存储在独立的NumPy数组中(数组长度均为33177600)。目前使用ASE库结合for循环生成XYZ文件,但循环效率极低,希望通过NumPy、Pandas等工具实现无循环的高效生成方案。

现有实现代码

import numpy as np
import ase
from ase import io
from ase import Atoms

atoms = []
for i in range(5):
    atoms.append(Atoms(
        ['Al','F','Al','F'], 
        positions=[
            (x_al1[i], y_al1[i], z_al1[i]), 
            (x_f1[i], y_f1[i], z_f1[i]),
            (x_al2[i], y_al2[i], z_al2[i]), 
            (x_f2[i], y_f2[i], z_f2[i])
        ]
    ))
ase.io.write(filename='tryj.xyz', images=atoms)

生成的XYZ文件示例(5种构型)

4
Properties=species:S:1:pos:R:3 pbc="F F F"
Al       2.57624552      -0.00195309       0.00000000
F        3.20148847       0.00277378       0.00000000
Al      -0.25834347       0.00195220      -0.00005894
F        0.36689949      -0.00277251       0.00008371
4
Properties=species:S:1:pos:R:3 pbc="F F F"
Al       2.57624552      -0.00195309       0.00000000
F        3.20148847       0.00277378       0.00000000
Al      -0.25834347       0.00188505      -0.00051102
F        0.36689949      -0.00267715       0.00072575
4
Properties=species:S:1:pos:R:3 pbc="F F F"
Al       2.57624552      -0.00195309       0.00000000
F        3.20148847       0.00277378       0.00000000
Al      -0.25834347       0.00149618      -0.00125539
F        0.36689949      -0.00212488       0.00178290
4
Properties=species:S:1:pos:R:3 pbc="F F F"
Al       2.57624552      -0.00195309       0.00000000
F        3.20148847       0.00277378       0.00000000
Al      -0.25834347       0.00058919      -0.00186210
F        0.36689949      -0.00083677       0.00264455
4
Properties=species:S:1:pos:R:3 pbc="F F F"
Al       2.57624552      -0.00195309       0.00000000
F        3.20148847       0.00277378       0.00000000
Al      -0.25834347      -0.00058919      -0.00186210
F        0.36689949       0.00083677       0.00264455

高效无循环实现方案(基于NumPy)

核心思路是利用NumPy向量化操作一次性构造所有文本内容,批量写入文件,彻底规避Python循环的性能损耗。

代码实现

import numpy as np

# 1. 合并原子坐标为统一数组
al1_coords = np.stack([x_al1, y_al1, z_al1], axis=1)
f1_coords = np.stack([x_f1, y_f1, z_f1], axis=1)
al2_coords = np.stack([x_al2, y_al2, z_al2], axis=1)
f2_coords = np.stack([x_f2, y_f2, z_f2], axis=1)
all_coords = np.stack([al1_coords, f1_coords, al2_coords, f2_coords], axis=1)

# 2. 定义原子种类和构型头部
species = np.array(['Al', 'F', 'Al', 'F'], dtype='U2')
header = "4\nProperties=species:S:1:pos:R:3 pbc=\"F F F\"\n"
headers = np.full(all_coords.shape[0], header, dtype='U100')

# 3. 向量化格式化所有坐标行
coord_lines = []
for i in range(4):
    # 对第i个原子的所有构型坐标生成格式化文本
    lines = np.char.add(
        species[i] + "       ",
        np.char.add(
            np.char.mod("%.8f      ", all_coords[:, i, 0]),
            np.char.add(
                np.char.mod("%.8f      ", all_coords[:, i, 1]),
                np.char.mod("%.8f", all_coords[:, i, 2])
            )
        )
    )
    coord_lines.append(lines)

# 4. 拼接每个构型的完整文本
configurations = np.char.add(
    np.char.add(coord_lines[0], "\n"),
    np.char.add(coord_lines[1], "\n"),
    np.char.add(coord_lines[2], "\n"),
    coord_lines[3]
)

# 5. 合并所有构型并写入文件
all_text = np.char.add(headers, configurations)
final_text = "\n".join(all_text) + "\n"

with open('all_configs.xyz', 'w') as f:
    f.write(final_text)

优化说明

  • 全程使用NumPy向量化运算,避免Python循环开销,处理3300多万组构型的速度比循环快数倍
  • 批量拼接文本后一次性写入,减少磁盘IO次数,进一步提升效率
  • 坐标格式化采用np.char系列函数实现向量字符串操作,无需逐元素处理

内容的提问来源于stack exchange,提问作者Mahmoud Amr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 04:33:10