如何用Python/Pandas实现文件列元素与另一文件对应值的除法运算
Got it, let's work through this problem step by step. We need to take the first column value from each line in file1, divide it by the corresponding numeric value for the same name in file2, then output the name and result to a new file. I'll show you two solid approaches: plain Python (no external dependencies) and Pandas (cleaner for tabular data workflows).
Approach 1: Plain Python (Lightweight, No Libraries)
This is great if you want to avoid installing extra packages and prefer a straightforward script. The core idea is to first load file2 into a dictionary for fast lookups, then process file1 line by line.
# First, load file2 into a dictionary: {name: numeric_value} name_to_value = {} with open('file2', 'r') as f2: for line in f2: # Split lines and skip any malformed entries parts = line.strip().split() if len(parts) != 2: continue name, value = parts[0], float(parts[1]) name_to_value[name] = value # Process file1 and write results to output with open('file1', 'r') as f1, open('output.txt', 'w') as out_file: for line in f1: parts = line.strip().split() # Skip lines that don't have the expected 3 columns if len(parts) != 3: continue file1_num = float(parts[0]) name = parts[2] # Avoid errors for missing names or division by zero if name in name_to_value: file2_num = name_to_value[name] if file2_num == 0: result = "NaN" # Customize this if you prefer skipping instead else: result = round(file1_num / file2_num, 4) out_file.write(f"{name} {result}\n")
Quick Notes:
- We add checks for malformed lines to prevent crashes from unexpected formatting.
- Division by zero is handled by writing "NaN"—adjust this to skip those lines or use a different placeholder if needed.
- The result is rounded to 4 decimal places to match your sample output; tweak the
round()argument for more/less precision.
Approach 2: Using Pandas (Cleaner for Tabular Data)
If you're working with larger datasets or already use Pandas in your projects, this method is more concise and readable.
First, install Pandas if you haven't already:
pip install pandas
Then run this script:
import pandas as pd # Read both files, assign column names since there are no headers df1 = pd.read_csv('file1', sep=r'\s+', names=['file1_num', 'line_label', 'name']) df2 = pd.read_csv('file2', sep=r'\s+', names=['name', 'file2_num']) # Merge the two datasets on the 'name' column (only keep matching entries) merged_data = pd.merge(df1[['name', 'file1_num']], df2, on='name', how='inner') # Calculate the division result and format it merged_data['result'] = merged_data['file1_num'] / merged_data['file2_num'] merged_data['result'] = merged_data['result'].round(4) # Write the final output without extra index/headers merged_data[['name', 'result']].to_csv('output_pandas.txt', sep=' ', index=False, header=False)
Quick Notes:
sep=r'\s+'handles any number of spaces/tabs as separators, which is useful if your files have inconsistent spacing.how='inner'ensures we only keep names present in both files. If you want to retain all names fromfile1(even those missing infile2), usehow='left'and fill missing values with something likemerged_data['file2_num'] = merged_data['file2_num'].fillna(1)(adjust based on your needs).
Both methods will produce output matching your example:
name1 0.2250 name2 0.1885 name3 0.0405
内容的提问来源于stack exchange,提问作者Steve Steveman Man

