You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas:如何基于现有列添加索引列,使重复项共享同一索引?

基于多列组合生成共享索引的解决方案

嘿,这个需求太常见啦!咱们用Pandas就能轻松搞定——核心就是根据old_index和year的组合给数据分组,让相同组合的行共享同一个索引,完全不用管num列的数值。下面给你两种简单好用的方法:

方法一:使用groupby().ngroup()

这是最直观的方式,先按目标列分组,再给每个组分配一个连续的索引值:

import pandas as pd

# 构造示例数据集
df = pd.DataFrame({
    'old_index': [1, 1, 2, 2, 3],
    'year': [2020, 2020, 2021, 2021, 2022],
    'num': [100, 200, 300, 400, 500]
})

# 生成新索引列
df['new_index'] = df.groupby(['old_index', 'year']).ngroup()

print(df)

输出结果:

old_index  year  num  new_index
0          1  2020  100          0
1          1  2020  200          0
2          2  2021  300          1
3          2  2021  400          1
4          3  2022  500          2

方法二:使用factorize()对组合列编码

如果处理大数据集追求更高效率,factorize()是个好选择,它直接对old_index和year的组合进行数值编码:

# 基于示例数据生成新索引
df['new_index'] = pd.factorize(df[['old_index', 'year']].apply(tuple, axis=1))[0]

print(df)

这个方法的输出和上面完全一致,编码逻辑更底层,处理大表时速度更快。

如果想要索引从1开始而非0,只需要在结果后加1即可,比如:df['new_index'] = df.groupby(['old_index', 'year']).ngroup() + 1。

内容的提问来源于stack exchange,提问作者Arist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 21:32:46