You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Pandas DataFrame中每个ID分配整数索引(附可复现示例)

如何为Pandas DataFrame设置每个ID对应的整数索引并保留ID列?

可复现的DataFrame示例

import pandas as pd
from io import StringIO

txt= """
ID,datetime,value
AB-CL-34,07/10/2022 10:00:00,5 
AB-CL-34,07/10/2022 11:15:10,7 
AB-CL-34,09/10/2022 15:30:30,13 
BX-RT-55,06/10/2022 11:30:22,0 
BX-RT-55,10/10/2022 22:44:11,1 
BX-RT-55,10/10/2022 23:30:22,6 
"""

df = pd.read_csv(StringIO(txt), parse_dates=[1], dayfirst=True)

需求说明

需要为每个唯一的ID分配一个连续整数作为索引,同时保留原有的ID列,最终输出格式如下:

ID            datetime value
0 AB-CL-34 07/10/2022 10:00:00     5 
0 AB-CL-34 07/10/2022 11:15:10     7 
0 AB-CL-34 09/10/2022 15:30:30    13 
1 BX-RT-55 06/10/2022 11:30:22     0 
1 BX-RT-55 10/10/2022 22:44:11     1 
1 BX-RT-55 10/10/2022 23:30:22     6 

解决方案

方法1:使用pd.factorize()

factorize()函数可直接将唯一的字符串ID映射为连续整数,将结果设置为索引即可:

df.index = pd.factorize(df['ID'])[0]

方法2:使用分类类型的cat.codes

将ID列转为category类型后,利用其cat.codes属性获取整数编码,再设置为索引:

df.index = df['ID'].astype('category').cat.codes

两种方法都能实现需求,factorize()更直接高效,category方式适合后续需要对ID进行分类操作的场景。


内容的提问来源于stack exchange,提问作者Johan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 22:40:16