You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中按另一列分组统计唯一ID数量并生成新列

问题:统计每个Reporting ID对应的唯一ID数量并生成新列

需求说明

统计数据集中每个Reporting ID对应的唯一ID数量,生成新列Reporting Count,规则为:

  • 仅当当前行的ID等于某个Reporting ID时,Reporting Count显示该Reporting ID对应的唯一ID总数
  • 其余行的Reporting Count显示0
  • 对于Reporting ID为空的行,Reporting Count显示该行ID作为Reporting ID时的计数

样本数据

ID     Reporting ID 
 101         103
 102         103
 103         107
 104         103 
 105         107 
 106         103 
 107         110
 108         103
 109         110 
 110          

期望输出

ID      Reporting ID     Reporting Count 
 101          103               0
 102          103               0
 103          107               5
 104          103               0
 105          107               0
 106          103               0
 107          110               2
 108          103               0
 109          110               0
 110                            2

尝试过的代码

df['Reporting Count'] = df.groupby('Reporting ID')['ID'].nunique()

问题分析

你当前的代码直接将分组统计结果赋值给新列,由于分组后的索引是Reporting ID,和原DataFrame的行索引不匹配,会导致大部分行的Reporting Count值为NaN,同时未实现需求中的条件显示逻辑。

解决方案

通过生成计数映射字典 + 条件赋值实现:

  1. 预处理数据并生成计数映射:
import pandas as pd
import numpy as np

# 将空字符串转为pd.NA,统一处理空值
df['Reporting ID'] = df['Reporting ID'].replace('', pd.NA)

# 统计每个Reporting ID对应的唯一ID数量,转为字典
count_map = df.groupby('Reporting ID')['ID'].nunique().to_dict()
# 补充空Reporting ID对应的计数:ID110作为Reporting ID时的计数是2
count_map[pd.NA] = count_map.get(110, 0)
  1. 条件生成目标列:
# 判断当前行ID是否是某个Reporting ID,是则取对应计数,否则设为0
df['Reporting Count'] = np.where(df['ID'].isin(count_map.keys()), df['ID'].map(count_map), 0)

运行后即可得到符合期望的输出。


内容的提问来源于stack exchange,提问作者Coding_Nubie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 05:22:30