You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas按关键词/分隔符起止解析势能分布(PED)数据

Clean Up & Parse Messy PED Files with Two Sections

Hey fellow chemist! I’ve wrestled with wonky PED output files (inconsistent rows, split sections) more times than I can count—so I totally get your frustration. Here’s a practical, Python-based solution that’ll split your file into the two sections you need and format everything into your 11 predefined columns.

Step 1: Parse the Main PED Section (Skip First 3 Rows, Stop at "***")

First, we’ll isolate the main data block by skipping the initial header rows and stopping at the "***" marker. We’ll also handle those annoying inconsistent row lengths by padding short rows with missing values and truncating overly long ones.

Code for Main Section

import pandas as pd

# Replace these with your exact 11 predefined column names
ped_column_names = [
    "tVib1", "%PED1", "tVib2", "%PED2", 
    "tVib3", "%PED3", "tVib4", "%PED4", 
    "tVib5", "%PED5", "Your11thColumn"  # Adjust this last one to match your needs
]

# Extract and clean the main PED data
main_ped_rows = []
with open("your_ped_file.txt", "r") as input_file:
    # Skip the first 3 rows
    for _ in range(3):
        next(input_file)
    
    # Read lines until we hit the "***" delimiter
    for line in input_file:
        stripped_line = line.strip()
        if stripped_line == "***":
            break
        if not stripped_line:  # Skip empty lines
            continue
        
        # Split line into values (handles multiple spaces/tabs)
        row_values = [val for val in stripped_line.split() if val]
        
        # Fix inconsistent row lengths
        while len(row_values) < len(ped_column_names):
            row_values.append(pd.NA)  # Pad with missing values if too short
        row_values = row_values[:len(ped_column_names)]  # Truncate if too long
        
        main_ped_rows.append(row_values)

# Convert to a clean DataFrame
main_ped_df = pd.DataFrame(main_ped_rows, columns=ped_column_names)
print("Cleaned Main PED Section:")
print(main_ped_df.head())

Step 2: Parse the "Alternative Coordinates" Section

Next, we’ll grab the second section starting right after the "alternative coordinates" header. We’ll use the same cleaning logic to ensure it fits your 11-column structure.

Code for Alternative Section

# Extract and clean the alternative coordinates data
alt_ped_rows = []
found_alt_header = False

with open("your_ped_file.txt", "r") as input_file:
    for line in input_file:
        stripped_line = line.strip()
        
        # Trigger data collection once we find the header
        if "alternative coordinates" in stripped_line.lower():
            found_alt_header = True
            continue  # Skip the header line itself
        
        if found_alt_header:
            if not stripped_line:
                continue
            
            # Same cleaning logic as the main section
            row_values = [val for val in stripped_line.split() if val]
            while len(row_values) < len(ped_column_names):
                row_values.append(pd.NA)
            row_values = row_values[:len(ped_column_names)]
            
            alt_ped_rows.append(row_values)

# Convert to DataFrame
alt_ped_df = pd.DataFrame(alt_ped_rows, columns=ped_column_names)
print("\nCleaned Alternative Coordinates Section:")
print(alt_ped_df.head())

Quick Tips for Edge Cases

  • Column name tweaks: Double-check that ped_column_names exactly matches your 11 predefined columns—swap out the placeholder names for your actual ones.
  • Odd formatting: If some rows have way more than 11 values, inspect the raw file to see if there’s extra metadata (like notes or footnotes) you need to exclude. You can add a check to skip lines that don’t match your expected value count.
  • Non-Python workflows: If you prefer command-line tools, you could use sed '/^$/d; 4,/^\\*\\*\\*/!d' your_ped_file.txt to extract the main section, then awk to format columns—but Python gives you more control for messy, inconsistent data.

内容的提问来源于stack exchange,提问作者HCSthe2nd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:27:47