如何基于event_id去除Python字典列表中的重复元素?
问题:字典列表按event_id去重失败
你有一个包含5个字典的列表,想按event_id字段去重(删除重复event_id的字典),但现有代码未生效,处理后列表长度仍为5。你的event_id值分别是"1"、"2"、"3"、"4"、"4",尝试的代码如下:
res = [] for y in self.dict_trie_csv_infos: #check if same 'event_id' before if float(y['event_id']) not in res: res.append(y)
错误原因
这段代码的问题很直接:你在判断float(y['event_id'])这个数值是否存在于res列表中,但res里存的是字典对象,不是单独的event_id数值,所以这个条件永远为True,所有字典都会被添加进去,自然无法实现去重。
解决方案
方案1:用集合跟踪已出现的event_id
通过一个集合记录已经处理过的event_id,遍历列表时只保留未出现过的字典:
res = [] seen_event_ids = set() for y in self.dict_trie_csv_infos: # 统一转换为float类型,避免字符串与数字的类型差异 event_id = float(y['event_id']) if event_id not in seen_event_ids: seen_event_ids.add(event_id) res.append(y)
这种方式会保留每个event_id第一次出现的字典。
方案2:用字典推导式快速去重
利用字典键的唯一性,以event_id为键实现去重:
# 此方法会保留每个event_id最后一次出现的字典 unique_dict = {float(y['event_id']): y for y in self.dict_trie_csv_infos} res = list(unique_dict.values())
如果想保留第一次出现的字典,只需反转遍历顺序再恢复:
unique_dict = {float(y['event_id']): y for y in reversed(self.dict_trie_csv_infos)} res = list(reversed(unique_dict.values()))
内容的提问来源于stack exchange,提问作者guiguilecodeur
相关产品推荐
相关产品推荐

