You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python基于变量观测值为每行生成唯一组合ID?

用Python生成多变量映射的唯一编码ID

实现思路

对每个变量的唯一观测值分配唯一数字(相同观测值对应同一数字),再将5个变量对应的数字按顺序拼接,得到每行的唯一编码。若两行所有变量值完全相同,编码也会一致。

代码实现(基于Pandas)

用Pandas的factorize()方法完成变量到数字的映射,该方法会自动给每个唯一值分配递增数字,再通过字符串拼接生成最终ID。

1. 准备示例数据

模拟匹配需求的示例数据集:

import pandas as pd

data = {
    "Var1": ["apple", "banana", "apple", "orange", "banana"],
    "Var2": ["red", "yellow", "red", "orange", "yellow"],
    "Var3": ["small", "large", "small", "medium", "large"],
    "Var4": ["round", "long", "round", "round", "long"],
    "Var5": ["sweet", "sweet", "sweet", "sour", "sweet"]
}
df = pd.DataFrame(data)

2. 为每个变量生成映射数字

对每个变量使用factorize(),并将结果加1(让数字从1开始,而非默认的0):

# 为每个变量生成编码列
for col in df.columns:
    df[f"{col}_code"] = pd.factorize(df[col])[0] + 1

3. 拼接编码生成唯一ID

将5个编码列的数值转为字符串后拼接,得到最终ID列:

# 拼接编码列生成唯一ID
df["Unique_ID"] = df[[f"{col}_code" for col in df.columns[:5]]].astype(str).agg(''.join, axis=1)

# 可选:删除中间生成的编码列
df = df.drop([f"{col}_code" for col in df.columns[:5]], axis=1)

4. 查看结果

打印处理后的数据集:

print(df)

输出结果符合需求逻辑:

Var1    Var2    Var3   Var4   Var5 Unique_ID
0   apple     red   small  round  sweet      11111
1  banana  yellow   long   long  sweet      22221
2   apple     red   small  round  sweet      11111
3  orange  orange  medium  round   sour      33322
4  banana  yellow   long   long  sweet      22221

可以看到,变量值完全相同的行(第0/2行、第1/4行),对应的Unique_ID也完全一致。

内容的提问来源于stack exchange,提问作者Paul Engelbert

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 23:40:22