如何基于字典的两个字段查找字典列表中的重复项?
问题描述
给定如下Python列表:
new_list = [ {'a':'2', 'b':'3', 'c':'1'}, {'a':'1', 'b':'2', 'c':'1'}, {'a':'1', 'b':'2', 'c':'1'}, {'a':'1', 'b':'4', 'c':'1'}, {'a':'2', 'b':'2', 'c':'1'}, {'a':'3', 'b':'2', 'c':'1'}]
需要找出仅基于a和b字段的重复项,预期结果为:
final_list = [{'a':'1', 'b':'2', 'c':'1'}]
尝试用集合解决,但作为Python新手遇到了困惑。
解决方案
字典是不可哈希类型,没法直接放进集合,所以得换个思路:先提取每个字典的a和b字段做成可哈希的元组,用这个元组来统计出现次数,再筛选出重复的项。下面给两种实用的实现方式:
方式一:手动用字典统计次数
new_list = [ {'a':'2', 'b':'3', 'c':'1'}, {'a':'1', 'b':'2', 'c':'1'}, {'a':'1', 'b':'2', 'c':'1'}, {'a':'1', 'b':'4', 'c':'1'}, {'a':'2', 'b':'2', 'c':'1'}, {'a':'3', 'b':'2', 'c':'1'}] # 统计每个(a,b)组合的出现次数 count_map = {} for item in new_list: key = (item['a'], item['b']) count_map[key] = count_map.get(key, 0) + 1 # 筛选重复项并去重 final_list = [] seen_keys = set() for item in new_list: key = (item['a'], item['b']) if count_map[key] > 1 and key not in seen_keys: final_list.append(item) seen_keys.add(key) print(final_list)
运行后就能得到预期结果:[{'a': '1', 'b': '2', 'c': '1'}]
方式二:用collections.Counter简化统计
如果不想自己写统计逻辑,可以用Python标准库的Counter工具,代码更简洁:
from collections import Counter new_list = [ {'a':'2', 'b':'3', 'c':'1'}, {'a':'1', 'b':'2', 'c':'1'}, {'a':'1', 'b':'2', 'c':'1'}, {'a':'1', 'b':'4', 'c':'1'}, {'a':'2', 'b':'2', 'c':'1'}, {'a':'3', 'b':'2', 'c':'1'}] # 生成所有(a,b)元组的计数结果 counter = Counter((item['a'], item['b']) for item in new_list) # 筛选重复项并去重 final_list = [] seen_keys = set() for item in new_list: key = (item['a'], item['b']) if counter[key] > 1 and key not in seen_keys: final_list.append(item) seen_keys.add(key) print(final_list)
这个逻辑和方式一完全一致,只是把统计部分交给Counter处理,代码更简洁易读。
内容的提问来源于stack exchange,提问作者user21274188
相关产品推荐
相关产品推荐

