You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas GroupBy根据分组内车型值生成责任车辆列?

问题描述

现有如下DataFrame:

accidentID      cartype   
0    58              70     
1    58              -70      
2    58              70     
3    58              100       
4    71              100   
5    71              -70    
6    250             70   
7    250             70  
8    250             100  
9    250             70  
10   70              70

需要新增一列car_in_the_wrong,规则优先级如下:

  • 若按accidentID分组内存在cartype = -70,则该组所有行的car_in_the_wrong为-70
  • 若分组内无-70但存在100,则该组所有行的car_in_the_wrong为100
  • 若-70和100都不存在,则该组所有行的car_in_the_wrong为70

期望得到的DataFrame如下:

accidentID      cartype        car_in_the_wrong
0    58              70             -70
1    58              -70            -70
2    58              70             -70
3    58              100            -70
4    71              100            -70
5    71              -70            -70
6    250             70             100
7    250             70             100
8    250             100            100
9    250             70             100
10   300             70             70

请问是否可以通过GroupBy实现该需求?


解决方案

完全可以通过GroupBy结合自定义逻辑实现,以下是两种高效的实现方式:

方法一:groupby.transform + 自定义函数

transform方法可以将分组计算的结果自动广播回原DataFrame的每一行,逻辑直观易懂:

import pandas as pd

# 构造原始DataFrame
df = pd.DataFrame({
    'accidentID': [58,58,58,58,71,71,250,250,250,250,70],
    'cartype': [70,-70,70,100,100,-70,70,70,100,70,70]
})

def get_wrong_car(group):
    if -70 in group['cartype'].values:
        return -70
    elif 100 in group['cartype'].values:
        return 100
    else:
        return 70

# 新增目标列
df['car_in_the_wrong'] = df.groupby('accidentID').transform(get_wrong_car)

方法二:groupby.agg + 映射

先计算每个分组的结果,再通过映射关联回原DataFrame,大数据量下性能更优:

# 计算每个accidentID对应的结果
group_map = df.groupby('accidentID')['cartype'].agg(
    lambda x: -70 if -70 in x.values else (100 if 100 in x.values else 70)
)

# 映射到原DataFrame
df['car_in_the_wrong'] = df['accidentID'].map(group_map)

两种方法都能得到符合预期的结果,可根据数据规模选择合适的方案。

内容的提问来源于stack exchange,提问作者Aleksander

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 12:50:45