Python Pandas入门:如何从反应式TXT构建化学计量矩阵?
Hey there! As someone new to Python and Pandas, let's break down exactly how to turn those reaction strings into the stoichiometric matrix you need—compounds as rows, reactions as columns. Here's a step-by-step guide with code you can follow:
Step 1: Parse Individual Reactions
First, we need a function to take a single reaction string (like R1: A + 2B + C <=> D) and extract the stoichiometric coefficients for each compound. Remember: reactants get negative coefficients (since they're consumed) and products get positive coefficients (since they're produced).
We'll use regular expressions to handle both compounds with explicit coefficients (like 2B) and those without (like A, which defaults to 1):
import re import pandas as pd def parse_reaction(reaction_str): # Split reaction name (e.g., R1) from the reaction equation rxn_name, rxn_eq = reaction_str.split(': ') # Split reactants (left) and products (right) reactants, products = rxn_eq.split(' <=> ') stoichiometry = {} # Process reactants (negative coefficients) for item in reactants.split(' + '): # Match optional number + compound name match = re.match(r'(\d*)([A-Za-z]+)', item) coeff = int(match.group(1)) if match.group(1) else 1 compound = match.group(2) stoichiometry[compound] = stoichiometry.get(compound, 0) - coeff # Process products (positive coefficients) for item in products.split(' + '): match = re.match(r'(\d*)([A-Za-z]+)', item) coeff = int(match.group(1)) if match.group(1) else 1 compound = match.group(2) stoichiometry[compound] = stoichiometry.get(compound, 0) + coeff return rxn_name, stoichiometry
Step 2: Read and Process the TXT File
Next, we'll read your reaction file line by line, use our parsing function on each line, and collect all the data we need:
# Initialize storage for reaction data and all unique compounds reaction_data = {} all_compounds = set() # Replace 'reactions.txt' with your actual file path with open('reactions.txt', 'r') as file: for line in file: line = line.strip() if not line: # Skip empty lines continue rxn_name, stoich = parse_reaction(line) reaction_data[rxn_name] = stoich # Add all compounds from this reaction to our set all_compounds.update(stoich.keys())
Step 3: Build the Stoichiometric Matrix DataFrame
Finally, we'll convert our collected data into a Pandas DataFrame. We'll sort the compound names for readability, fill missing values (compounds not in a reaction) with 0, and convert to integers:
# Create DataFrame, sort rows (compounds) alphabetically stoichiometric_matrix = pd.DataFrame( reaction_data, index=sorted(all_compounds) ).fillna(0).astype(int) # Print the result print(stoichiometric_matrix)
Example Output
Using your sample reactions:
R1: A + 2B + C <=> D
R2: A + B <=> C
You'll get this matrix:
R1 R2 A -1 -1 B -2 -1 C -1 1 D 1 0
This matches exactly what you're looking for: each row is a compound, each column is a reaction, and the values are the stoichiometric coefficients.
内容的提问来源于stack exchange,提问作者Microdot

