如何在Pandas中添加带前缀的唯一标识符列
Got it, let's break down how to add that custom unique ID column with the "ACC-" prefix based on the City values. Here are a few reliable approaches depending on your needs:
方法1:自动按City出现顺序分配ID(简洁版)
If you just need each unique city to get a sequential number starting from 1 (same cities share the same ID regardless of their order), use groupby().ngroup():
import pandas as pd # 示例DataFrame df = pd.DataFrame({'City': ['Atlanta', 'Newyork', 'Atlanta', 'Chicago', 'Newyork']}) # 直接生成目标列 df['ACC_ID'] = 'ACC-' + (df.groupby('City').ngroup() + 1).astype(str)
ngroup() assigns a unique integer to each group starting from 0, so we add 1 to make it start at 1, then convert to string and prepend "ACC-".
方法2:用factorize显式生成编码
Another common method is pd.factorize(), which maps unique values to integer codes. This works similarly but lets you see the intermediate ID if needed:
# 生成City对应的数字ID(从1开始) df['City_Code'] = pd.factorize(df['City'])[0] + 1 # 拼接前缀得到最终列 df['ACC_ID'] = 'ACC-' + df['City_Code'].astype(str)
方法3:手动指定City对应的ID(自定义映射)
If you need to enforce specific numbers for certain cities (e.g., Atlanta must be ACC-1, Newyork must be ACC-2 no matter their order in the DataFrame), create a custom mapping dictionary:
# 自定义City到ID的映射 city_to_id = {'Atlanta': 1, 'Newyork': 2, 'Chicago': 3} # 应用映射并拼接前缀 df['ACC_ID'] = 'ACC-' + df['City'].map(city_to_id).astype(str)
最终效果示例
After running any of the above, your DataFrame will look like this:
| City | ACC_ID |
|---|---|
| Atlanta | ACC-1 |
| Newyork | ACC-2 |
| Atlanta | ACC-1 |
| Chicago | ACC-3 |
| Newyork | ACC-2 |
内容的提问来源于stack exchange,提问作者Daven1

