如何为Pandas DataFrame的分组id添加重置式递增索引列?
Pandas分组内生成递增序列的最优实现
你要的需求用Pandas内置的groupby+cumcount()就能高效实现,完全不用手动遍历唯一id,这是最简洁且性能最优的方案:
示例代码
import pandas as pd # 构造原始DataFrame df = pd.DataFrame({ 'id': [0, 0, 1, 1, 1, 1, 2, 2, 2], 'animal': ['dog', 'cat', 'goat', 'cow', 'sheep', 'pig', 'lion', 'tiger', 'bear'] }) # 添加ix列:按id分组后,每组内从0开始计数 df['ix'] = df.groupby('id').cumcount() print(df)
输出结果
id animal ix 0 0 dog 0 1 0 cat 1 2 1 goat 0 3 1 cow 1 4 1 sheep 2 5 1 pig 3 6 2 lion 0 7 2 tiger 1 8 2 bear 2
说明
cumcount()方法会自动对每个分组内的行从0开始顺序计数,遇到新分组时自动重置为0,完全匹配你的需求。相比手动遍历唯一id的方法,它是Pandas底层优化的向量化操作,在数据量较大时性能优势非常明显,代码也更简洁易读。
内容的提问来源于stack exchange,提问作者M. Fire
相关产品推荐
相关产品推荐

