如何为pandas DataFrame中两列的无序组合独立分配唯一ID
实现方法
核心逻辑是先把每行的col1和col2按大小排序,消除顺序差异,得到唯一的组合标识,再对组合标识编码生成ID。
方法1:逐行处理(小数据量适用)
代码如下:
import pandas as pd # 构造示例DataFrame df = pd.DataFrame({ 'col1': [1, 2, 2, 3, 3, 4], 'col2': [2, 1, 3, 2, 4, 3] }) # 生成无序组合的唯一ID df['id'] = df.apply(lambda row: tuple(sorted([row['col1'], row['col2']])), axis=1).factorize()[0] + 1 print(df)
输出结果和要求完全一致:
col1 col2 id 0 1 2 1 1 2 1 1 2 2 3 2 3 3 2 2 4 3 4 3 5 4 3 3
方法2:向量化操作(大数据量适用)
如果DataFrame行数较多,apply逐行处理性能偏低,可以用numpy向量化排序提升效率:
import pandas as pd import numpy as np df = pd.DataFrame({ 'col1': [1, 2, 2, 3, 3, 4], 'col2': [2, 1, 3, 2, 4, 3] }) # 向量化排序后编码生成ID sorted_arr = np.sort(df[['col1', 'col2']].values, axis=1) df['id'] = pd.Series(list(zip(sorted_arr[:, 0], sorted_arr[:, 1]))).factorize()[0] + 1
内容的提问来源于stack exchange,提问作者tizipupi
相关产品推荐
相关产品推荐

