查找DataFrame跨组重复值并获取对应关联列值的技术需求
Solution for Matching Values and Mapping Associated Columns in Pandas
Got it, let's break this down step by step to get your desired res and res1 columns. Here's a straightforward approach using pandas:
Step 1: Set up your DataFrame
First, let's recreate your sample data to work with:
import pandas as pd data = { 'A': [44.562, 42.621, 38.751, 44.562, 40.472], 'B': [42.975, 40.79, 39.301, 41.229, 41.278], 'C': [39.252, 40.741, 43.756, 43.2, 43.951], 'D': [40.973, 40.887, 40.68, 41.59, 41.99], 'R': [37.994, 40.863, 38.397, 43.561, 42.55], 'S': [39.191, 40.204, 39.301, 41.01, 41.083], 'T': [38.69, 40.131, 42.133, 43.2, 42.572], 'U': [40.973, 40.887, 38.788, 41.486, 40.472] } df = pd.DataFrame(data)
Step 2: Create the res column
We need to find the value in the R/S/T/U group that exists in the A/B/C/D group for each row. We can use apply() to iterate over each row and find this matching value:
# Define the two column groups group1 = ['A', 'B', 'C', 'D'] group2 = ['R', 'S', 'T', 'U'] # Extract the matching value for res df['res'] = df.apply(lambda row: next(val for val in row[group2] if val in row[group1]), axis=1)
Step 3: Create the res1 column using the association rules
Next, we map the matched column from group1 to its associated column (A ↔ D, B ↔ C) and pull the corresponding value:
# Define the association mapping assoc_map = {'A': 'D', 'D': 'A', 'B': 'C', 'C': 'B'} def get_res1(row): # Find which column in group1 matches the res value matched_col = next(col for col in group1 if row[col] == row['res']) # Return the value from the associated column return row[assoc_map[matched_col]] df['res1'] = df.apply(get_res1, axis=1)
Final Result
If you print df now, you'll get exactly the output you expected:
A B C D R S T U res res1 0 44.562 42.975 39.252 40.973 37.994 39.191 38.69 40.973 40.973 44.562 1 42.621 40.790 40.741 40.887 40.863 40.204 40.131 40.887 40.887 42.621 2 38.751 39.301 43.756 40.680 38.397 39.301 42.133 38.788 39.301 43.756 3 44.562 41.229 43.200 41.590 43.561 41.010 43.200 41.486 43.200 41.229 4 40.472 41.278 43.951 41.990 42.550 41.083 42.572 40.472 40.472 41.990
Notes
- This approach assumes exactly one matching value per row between the two groups (which aligns with your problem statement).
- The
next()function works here because we know there's always a match—if there were cases with no matches, you'd want to add error handling (like a default value).
内容的提问来源于stack exchange,提问作者moys
相关产品推荐
相关产品推荐

