如何为特定结构DataFrame新增列及计算不同组点的距离?
Hey there! Let's break down your two pandas and distance-calculation questions with practical, efficient solutions.
First, let's ground this in a realistic scenario: your DataFrame has one column with repeated values (like A_id, tied to multiple entries per unique B_id) and another column with unique identifiers (like B_id). We can use pandas' built-in tools to add a new column without slow loops.
Example DataFrame
Suppose your data looks like this:
import pandas as pd data = { 'B_id': [123, 123, 456, 456], 'A_id': [1, 2, 3, 4], 'score': [85, 90, 78, 82], 're-score': [88, 92, 75, 80], 'class': [1, 0, 1, 0] # Your "flag=1" category } df = pd.DataFrame(data)
Common Use Case: Add a Group Reference Value
If you want to add a column that pulls the score from the class=1 entry in each B_id group, here's an efficient way:
# Get the reference score for each B_id where class=1 ref_scores = df[df['class'] == 1].set_index('B_id')['score'] # Map this reference value to every row in its B_id group df['ref_score'] = df['B_id'].map(ref_scores)
Quick Tips
- Use
groupby().transform()for group-level stats (like mean, max) that need to repeat for every row in the group. - Skip
iterrows()for column creation—it's slow for large datasets.
Your current code uses a single reference row, but let's adjust it to handle grouping properly (since each B_id has multiple A_id entries).
Case 1: Distance to Group's Class=1 Reference
If you want to compute the distance from every row in a B_id group to the group's own class=1 entry, use groupby() to process each group independently:
from scipy.spatial.distance import euclidean def process_group(group): # Grab the class=1 row as the reference (assumes one per group) ref_row = group[group['class'] == 1].iloc[0] # Define columns to use for distance calculation distance_cols = ['score', 're-score'] # Calculate distance for each row in the group group['distance_to_ref'] = group[distance_cols].apply( lambda row: euclidean(row, ref_row[distance_cols]), axis=1 ) return group # Apply the function to each B_id group df_with_distances = df.groupby('B_id').apply(process_group).reset_index(drop=True)
Case 2: Cross-Group Pairwise Distances
If you need to calculate distances between every point in one B_id group and every point in another, use scipy.spatial.distance.cdist for speed:
from scipy.spatial.distance import cdist # Extract point matrices for each B_id group group_points = { bid: group[['score', 're-score']].values for bid, group in df.groupby('B_id') } # Calculate cross-group pairwise distances (avoid duplicate pairs) cross_group_distances = [] for bid1, points1 in group_points.items(): for bid2, points2 in group_points.items(): if bid1 < bid2: dist_matrix = cdist(points1, points2, metric='euclidean') cross_group_distances.append({ 'group1': bid1, 'group2': bid2, 'distance_matrix': dist_matrix.tolist() }) # Convert results to a readable DataFrame distance_df = pd.DataFrame(cross_group_distances)
Key Improvements Over Your Original Code
- Each
B_idgroup uses its ownclass=1reference, not a single global row. cdistandgroupbyare way more efficient thaniterrows()for large datasets.
内容的提问来源于stack exchange,提问作者Nari Kumar

