You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为特定结构DataFrame新增列及计算不同组点的距离?

Hey there! Let's break down your two pandas and distance-calculation questions with practical, efficient solutions.

1. Adding a New Column to Your Multi-Value + Unique-Value DataFrame

First, let's ground this in a realistic scenario: your DataFrame has one column with repeated values (like A_id, tied to multiple entries per unique B_id) and another column with unique identifiers (like B_id). We can use pandas' built-in tools to add a new column without slow loops.

Example DataFrame

Suppose your data looks like this:

import pandas as pd
data = {
    'B_id': [123, 123, 456, 456],
    'A_id': [1, 2, 3, 4],
    'score': [85, 90, 78, 82],
    're-score': [88, 92, 75, 80],
    'class': [1, 0, 1, 0]  # Your "flag=1" category
}
df = pd.DataFrame(data)

Common Use Case: Add a Group Reference Value

If you want to add a column that pulls the score from the class=1 entry in each B_id group, here's an efficient way:

# Get the reference score for each B_id where class=1
ref_scores = df[df['class'] == 1].set_index('B_id')['score']
# Map this reference value to every row in its B_id group
df['ref_score'] = df['B_id'].map(ref_scores)

Quick Tips

  • Use groupby().transform() for group-level stats (like mean, max) that need to repeat for every row in the group.
  • Skip iterrows() for column creation—it's slow for large datasets.
2. Calculating Distances Between Points in Groups

Your current code uses a single reference row, but let's adjust it to handle grouping properly (since each B_id has multiple A_id entries).

Case 1: Distance to Group's Class=1 Reference

If you want to compute the distance from every row in a B_id group to the group's own class=1 entry, use groupby() to process each group independently:

from scipy.spatial.distance import euclidean

def process_group(group):
    # Grab the class=1 row as the reference (assumes one per group)
    ref_row = group[group['class'] == 1].iloc[0]
    # Define columns to use for distance calculation
    distance_cols = ['score', 're-score']
    # Calculate distance for each row in the group
    group['distance_to_ref'] = group[distance_cols].apply(
        lambda row: euclidean(row, ref_row[distance_cols]),
        axis=1
    )
    return group

# Apply the function to each B_id group
df_with_distances = df.groupby('B_id').apply(process_group).reset_index(drop=True)

Case 2: Cross-Group Pairwise Distances

If you need to calculate distances between every point in one B_id group and every point in another, use scipy.spatial.distance.cdist for speed:

from scipy.spatial.distance import cdist

# Extract point matrices for each B_id group
group_points = {
    bid: group[['score', 're-score']].values 
    for bid, group in df.groupby('B_id')
}

# Calculate cross-group pairwise distances (avoid duplicate pairs)
cross_group_distances = []
for bid1, points1 in group_points.items():
    for bid2, points2 in group_points.items():
        if bid1 < bid2:
            dist_matrix = cdist(points1, points2, metric='euclidean')
            cross_group_distances.append({
                'group1': bid1,
                'group2': bid2,
                'distance_matrix': dist_matrix.tolist()
            })

# Convert results to a readable DataFrame
distance_df = pd.DataFrame(cross_group_distances)

Key Improvements Over Your Original Code

  • Each B_id group uses its own class=1 reference, not a single global row.
  • cdist and groupby are way more efficient than iterrows() for large datasets.

内容的提问来源于stack exchange,提问作者Nari Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:09:22