使用Pandas按关键词/分隔符起止解析势能分布(PED)数据
Hey fellow chemist! I’ve wrestled with wonky PED output files (inconsistent rows, split sections) more times than I can count—so I totally get your frustration. Here’s a practical, Python-based solution that’ll split your file into the two sections you need and format everything into your 11 predefined columns.
Step 1: Parse the Main PED Section (Skip First 3 Rows, Stop at "***")
First, we’ll isolate the main data block by skipping the initial header rows and stopping at the "***" marker. We’ll also handle those annoying inconsistent row lengths by padding short rows with missing values and truncating overly long ones.
Code for Main Section
import pandas as pd # Replace these with your exact 11 predefined column names ped_column_names = [ "tVib1", "%PED1", "tVib2", "%PED2", "tVib3", "%PED3", "tVib4", "%PED4", "tVib5", "%PED5", "Your11thColumn" # Adjust this last one to match your needs ] # Extract and clean the main PED data main_ped_rows = [] with open("your_ped_file.txt", "r") as input_file: # Skip the first 3 rows for _ in range(3): next(input_file) # Read lines until we hit the "***" delimiter for line in input_file: stripped_line = line.strip() if stripped_line == "***": break if not stripped_line: # Skip empty lines continue # Split line into values (handles multiple spaces/tabs) row_values = [val for val in stripped_line.split() if val] # Fix inconsistent row lengths while len(row_values) < len(ped_column_names): row_values.append(pd.NA) # Pad with missing values if too short row_values = row_values[:len(ped_column_names)] # Truncate if too long main_ped_rows.append(row_values) # Convert to a clean DataFrame main_ped_df = pd.DataFrame(main_ped_rows, columns=ped_column_names) print("Cleaned Main PED Section:") print(main_ped_df.head())
Step 2: Parse the "Alternative Coordinates" Section
Next, we’ll grab the second section starting right after the "alternative coordinates" header. We’ll use the same cleaning logic to ensure it fits your 11-column structure.
Code for Alternative Section
# Extract and clean the alternative coordinates data alt_ped_rows = [] found_alt_header = False with open("your_ped_file.txt", "r") as input_file: for line in input_file: stripped_line = line.strip() # Trigger data collection once we find the header if "alternative coordinates" in stripped_line.lower(): found_alt_header = True continue # Skip the header line itself if found_alt_header: if not stripped_line: continue # Same cleaning logic as the main section row_values = [val for val in stripped_line.split() if val] while len(row_values) < len(ped_column_names): row_values.append(pd.NA) row_values = row_values[:len(ped_column_names)] alt_ped_rows.append(row_values) # Convert to DataFrame alt_ped_df = pd.DataFrame(alt_ped_rows, columns=ped_column_names) print("\nCleaned Alternative Coordinates Section:") print(alt_ped_df.head())
Quick Tips for Edge Cases
- Column name tweaks: Double-check that
ped_column_namesexactly matches your 11 predefined columns—swap out the placeholder names for your actual ones. - Odd formatting: If some rows have way more than 11 values, inspect the raw file to see if there’s extra metadata (like notes or footnotes) you need to exclude. You can add a check to skip lines that don’t match your expected value count.
- Non-Python workflows: If you prefer command-line tools, you could use
sed '/^$/d; 4,/^\\*\\*\\*/!d' your_ped_file.txtto extract the main section, thenawkto format columns—but Python gives you more control for messy, inconsistent data.
内容的提问来源于stack exchange,提问作者HCSthe2nd

