You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除包含DataFrame的元组列表中的重复项

问题描述

需要移除包含数字和Pandas DataFrame的元组列表中的重复项,只有当元组中的数字和DataFrame完全相同时,才判定为重复项。

示例输入

[  ( 5.42,
      c1     c2    c3     c4     
r1    1      1      1      1    
r2    1      0      0      0   
r3    0      1      0      1     
r4    0      0      0      0 ),

  ( 4.31,
      c1    c2     c3     c4     
r1    0     1      0      1    
r2    0     1      1      1    
r3    0     0     0     0   
r4    1     0     0     1 ),

  ( 5.42,
      c1     c2    c3    c4     
r1    0      0      0     0    
r2    0      0      1     0     
r3    0      0      0     1   
r4    1      1      1     1 ),

  ( 5.42,
      c1    c2    c3    c4     
r1    0     0     0     0    
r2    0     0      1     0     
r3    0     0     0     1   
r4    1      1      1     1 )  ]

期望输出

[  ( 5.42,
      c1    c2    c3    c4     
r1    1     1     1     1    
r2    1     0     0     0   
r3    0     1     0     1     
r4    0     0     0     0),

  ( 4.31,
      c1    c2    c3    c4     
r1    0     1     0     1    
r2    0     1     1     1    
r3    0     0     0     0   
r4    1     0     0     1 ),

  ( 5.42,
      c1    c2    c3    c4     
r1    0     0     0     0    
r2    0     0     1     0     
r3    0     0     0     1   
r4    1      1      1     1 )  ]

遇到的错误

尝试多种方法后,总是触发以下两类错误:

  • ValueError: The truth value of the DataFrame is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all()
  • TypeError: DataFrame objects are mutable, thus cannot be hashed

尝试过的代码

  • 尝试1:
    from itertools import groupby
    deduped = [next(g) for k, g in groupby(sorted(tup_list), key=lambda tup: tup[:2])]
    
  • 尝试2:
    new_tup_list = [tuple(item) for item in tup_list]
    deduped = list(set(new_tup_list))
    
  • 尝试3:
    deduped = list(set(map(tuple, tup_list)))
    
  • 尝试4:
    from collections import OrderedDict
    d = OrderedDict()
    for t in tup_list:
        d.setdefault(t[1], t)
    deduped = list(d.values())
    
  • 尝试5:
    deduped = set()
    temp = [deduped.add((a,b)) for (a,b) in tup_list if (a,b) not in deduped]
    
解决方案

核心问题是DataFrame不可哈希,且直接比较会触发歧义错误,所以需要把DataFrame转换成可哈希、可比较的类型,再进行去重操作。

方法1:将DataFrame转为元组的元组(最严谨)

把DataFrame的索引、列名、数据全部转换成不可变的元组,确保只有完全匹配时才视为相同:

def df_to_hashable(df):
    # 把索引、列、数据都转成元组,保留完整结构信息
    return (tuple(df.index), tuple(df.columns), tuple(tuple(row) for row in df.values))

seen = set()
deduped = []
for num, df in tup_list:
    # 生成当前元组的唯一标识
    hash_key = (num, df_to_hashable(df))
    if hash_key not in seen:
        seen.add(hash_key)
        deduped.append((num, df))

方法2:使用DataFrame哈希值(效率更高)

通过Pandas内置工具生成DataFrame的哈希值,结合数字作为判断重复的依据:

import pandas as pd

seen = set()
deduped = []
for num, df in tup_list:
    # 生成包含索引和列的哈希值,确保结构+数据完全匹配
    df_hash = tuple(pd.util.hash_pandas_object(df, index=True, columns=True))
    hash_key = (num, df_hash)
    if hash_key not in seen:
        seen.add(hash_key)
        deduped.append((num, df))

方法3:使用groupby保留原顺序

如果需要保留原列表中首次出现的顺序,可先按自定义键排序后再分组去重:

from itertools import groupby

def get_key(tup):
    num, df = tup
    return (num, tuple(df.index), tuple(df.columns), tuple(tuple(row) for row in df.values))

# 先按自定义键排序,确保相同项相邻
sorted_list = sorted(tup_list, key=get_key)
deduped = [next(g) for k, g in groupby(sorted_list, key=get_key)]

说明

  • 方法1是最严谨的,完全匹配元组中的数字、DataFrame的索引、列名和所有数据值
  • 方法2效率更高,但存在极低的哈希碰撞概率
  • 所有方法都避开了直接操作DataFrame的哈希和比较,解决了报错问题

内容的提问来源于stack exchange,提问作者user23487612

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 07:00:55