You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas多列去重:两列交换位置后值相同视为重复的实现方案

Pandas实现无序特征对去重方法

核心思路

要解决特征对顺序不敏感的去重需求,本质是把顺序相关的feature1、feature2组合,转换为顺序无关的唯一标识,只要两个特征的元素完全相同,不管顺序如何,标识值都一致,再基于标识去重即可。

完整实现代码

import pandas as pd

# 1. 构造示例数据集
df = pd.DataFrame({
    'feature1': ['a', 'b', 'a', 'a', 'a'],
    'feature2': ['b', 'a', 'c', 'd', 'd'],
    'pcc': [0.6, 0.4, 0.7, -0.1, 0.3]
})

# 2. 生成无序对唯一标识:将两个特征排序后转为元组
df['pair_key'] = df.apply(lambda x: tuple(sorted([x['feature1'], x['feature2']])), axis=1)

# 3. 按标识去重,保留首次出现的行,删除辅助列
df_result = df.drop_duplicates(subset='pair_key', keep='first').drop(columns='pair_key')

print(df_result)

运行后输出结果和预期完全一致:

feature1 feature2  pcc
0        a        b  0.6
2        a        c  0.7
3        a        d -0.1

大数据量优化方案

如果数据集行数超过10万,apply逐行操作性能较低,可以改用numpy向量化排序提升效率:

import numpy as np
# 向量化生成无序对标识
df['pair_key'] = [tuple(row) for row in np.sort(df[['feature1', 'feature2']].values, axis=1)]
df_result = df.drop_duplicates('pair_key', keep='first').drop('pair_key', axis=1)

内容的提问来源于stack exchange,提问作者ah bon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 11:18:05