如何用Pandas提取子集与超集CSV的交集列及超集指定列?
Here's a straightforward, efficient way to get the exact result you're looking for using Pandas:
Step 1: Set Up Pandas
First, make sure Pandas is installed (run pip install pandas if you haven't already), then import it:
import pandas as pd
Step 2: Load Your CSV Files
Read both your super.csv (the full dataset) and sub.csv (the subset of Num values) into Pandas DataFrames:
# Load the super set with Num and Val columns df_super = pd.read_csv('super.csv') # Load the subset with only Num column df_sub = pd.read_csv('sub.csv')
Step 3: Filter for Matching Rows
Use Pandas' built-in isin() method to filter the super DataFrame down to only rows where Num exists in the subset:
# Extract the list of Num values we want to keep target_nums = df_sub['Num'].tolist() # Filter the super DataFrame to retain only matching rows result_df = df_super[df_super['Num'].isin(target_nums)]
Step 4: Output the Result
Print the data in your requested format, or save it to a new CSV if needed:
# Print the result with the exact format you specified print("Num , Val") for _, row in result_df.iterrows(): print(f"{row['Num']},{row['Val']}") # Optional: Save to a new CSV file result_df.to_csv('intersection_result.csv', index=False)
Example Output
When you run the code, you'll get:
Num , Val 2,25 4,87
This method leverages Pandas' optimized filtering capabilities, so it works smoothly even if your datasets grow larger over time.
内容的提问来源于stack exchange,提问作者Vatsalya Tandon

