You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python pandas结合groupby与np.where实现R data.table分组条件赋值

解决方案

核心思路是使用pandas的groupby.transform()方法,它会将分组聚合后的结果映射回原数据的每一行,完美匹配你需要的按id分组取unitsnew最小值的需求。

完整可运行代码如下:

import pandas as pd
import numpy as np

matrix = [(1, 34, 23),
          (2, 31, 11),
          (3, 16, 21),
          (4, 32, 22),
          (1, np.nan, 27),
          (5, 35, 11)]
df = pd.DataFrame(matrix, columns = ['id', 'units', 'unitsnew'])

# 按id分组计算unitsnew的分组最小值,再做条件替换
df['col'] = np.where(
    df['units'].isna(),
    df.groupby('id')['unitsnew'].transform(np.nanmin),
    df['units']
)

print(df)

运行输出结果:

id  units  unitsnew   col
0   1   34.0        23  34.0
1   2   31.0        11  31.0
2   3   16.0        21  16.0
3   4   32.0        22  32.0
4   1    NaN        27  23.0
5   5   35.0        11  35.0

如果你偏好链式调用的assign写法,也可以用下面的实现:

df = df.assign(
    col = lambda x: np.where(
        x['units'].isna(),
        x.groupby('id')['unitsnew'].transform(np.nanmin),
        x['units']
    )
)

内容的提问来源于stack exchange,提问作者Mike

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 12:54:00