如何合并DataFrame中列值反转的重复行?解决iterrows报错问题
合并DataFrame中列值互为反转的重复行及报错解决
问题需求
现有如下结构的DataFrame:
Column1 Column2 A B B A C D D C E F
需要合并其中列值互为反转的重复行,最终得到:
Column1 Column2 A B C D E F
需处理1000个行数不足50行的小文件,原尝试使用iterrows编写代码但报错。
报错原因及修复
直接报错修复
原代码中row_rev_index = df[(df['Column1'] == row['Column2']) & (df['Column2'] == row['Column1'])].index()触发TypeError: 'Int64Index' object is not callable,原因是**index是DataFrame的属性而非可调用方法**,不能加括号,应改为.index。
但原代码还存在逻辑缺陷:比如循环内每次重置output = []会丢失之前的结果,且会重复处理同一对反转行,以下提供两种更合理的解决方案。
高效向量化解决方案(推荐)
利用向量化操作对每行的两列值排序,再基于排序结果去重,适合批量处理大量小文件,效率远高于iterrows:
import pandas as pd import numpy as np # 处理单个文件的示例 df = pd.read_csv("your_file_path.csv") # 对每行的Column1和Column2进行排序,生成临时列 df[['sorted_col1', 'sorted_col2']] = pd.DataFrame( np.sort(df[['Column1', 'Column2']], axis=1), index=df.index ) # 基于临时列去重,保留每组反转行的第一行,再删除临时列 df_unique = df.drop_duplicates(subset=['sorted_col1', 'sorted_col2']).drop(columns=['sorted_col1', 'sorted_col2']) print(df_unique)
修复后的iterrows版本
如果坚持使用遍历方式,需添加已处理配对的记录,避免重复操作:
import pandas as pd df = pd.read_csv("your_file_path.csv") output = [] seen_pairs = set() # 记录已处理的配对,避免重复 for index, row in df.iterrows(): # 将配对转为有序元组,统一A-B和B-A的判断标准 current_pair = tuple(sorted((row['Column1'], row['Column2']))) if current_pair not in seen_pairs: seen_pairs.add(current_pair) output.append(row) # 转为最终的DataFrame df_unique = pd.DataFrame(output) print(df_unique)
内容的提问来源于stack exchange,提问作者zzz
相关产品推荐
相关产品推荐

