You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从.out文件提取数据构建指定三列的DataFrame

提取.out文件数据并生成指定DataFrame的解决方案

问题描述

需要用Python从.out文件提取数据,生成包含三列的DataFrame:

  • 分子名称:对应文件名
  • Zero-point correction:提取对应数值
  • Sum of electronic and zero-point Energies:提取对应数值

现有代码已能定位目标数据,但无法将数据整理到DataFrame中,原代码及输出示例如下:

原代码

import os
import pandas as pd
files = [f for f in os.listdir(".") if '.out' in f]
words = [' Zero-point correction', 'Sum of electronic and zero-point Energies']


for f in files:
    with open(f, 'r') as file:
          lines = file.readlines()
          for line in lines:
            for word in words:  
             if word in line:
                  a,b = line.split('=')
                 
                  print(a,b,f)

原输出示例

Zero-point correction                            0.171551 (Hartree/Particle)
 O-1.out
 Sum of electronic and zero-point Energies            -738.826005
 O-1.out

解决方案

修改代码,通过列表收集每个文件的提取结果,再转换为DataFrame:

import os
import pandas as pd

# 获取当前目录下所有.out文件
files = [f for f in os.listdir(".") if '.out' in f]
# 定义需要提取的目标关键词
target_keys = [
    ' Zero-point correction', 
    'Sum of electronic and zero-point Energies'
]

# 初始化列表,存储每个文件的提取数据
data_list = []

for filename in files:
    # 初始化当前文件的结果字典,确保结构统一
    file_data = {
        '分子名称': filename,
        'Zero-point correction': None,
        'Sum of electronic and zero-point Energies': None
    }
    
    with open(filename, 'r') as file:
        lines = file.readlines()
        for line in lines:
            # 匹配Zero-point correction行
            if target_keys[0] in line:
                _, value_part = line.split('=')
                # 去除括号内容和多余空格,转换为数值
                zpe_value = float(value_part.split('(')[0].strip())
                file_data['Zero-point correction'] = zpe_value
            # 匹配Sum of electronic...行
            elif target_keys[1] in line:
                _, value_part = line.split('=')
                sum_energy_value = float(value_part.strip())
                file_data['Sum of electronic and zero-point Energies'] = sum_energy_value
    
    # 将当前文件的结果加入列表
    data_list.append(file_data)

# 转换为DataFrame
result_df = pd.DataFrame(data_list)
print(result_df)

代码说明

  • 用data_list存储每个文件的提取结果,每个元素是包含三列数据的字典,确保数据结构统一
  • 针对Zero-point correction行,额外处理了括号内的单位内容,只保留纯数值
  • 即使某文件缺失某条数据,字典的None值也能保证DataFrame结构完整
  • 最后通过pd.DataFrame()直接将列表转换为标准结构的DataFrame

示例输出

分子名称  Zero-point correction  Sum of electronic and zero-point Energies
0  O-1.out                0.171551                                -738.826005

内容的提问来源于stack exchange,提问作者AIme999

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 04:55:23