You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于指定键(name、idVendor、idProduct)去除字典列表重复项?

基于指定键去除字典列表中的重复项

给定如下字典列表:

list_ = [{'printer': {'name': 'GEZHI, micro-printer', 'type': 'Usb', 'idVendor': 10473, 'idProduct': 649, 'in_ep': 129, 'out_ep': 1, 'filename': 'config1.yaml', 'status': 1}}, 
         {'printer': {'name': 'TECH, CLA58', 'type': 'Usb', 'idVendor': 26728, 'idProduct': 512, 'in_ep': 130, 'out_ep': 1, 'filename': 'config2.yaml', 'status': 0}}, 
         {'printer': {'name': 'GEZHI, micro-printer', 'type': 'Usb', 'idVendor': '10473', 'idProduct': '649', 'status': 2, 'out_ep': '1', 'in_ep': '129'}}]

需求是:仅基于name、idVendor和idProduct这三个键去除列表中的重复项。

你尝试的代码及错误

你之前尝试基于name去重的代码:

res = [dict(t['printer']['name']) for t in {tuple(d.items()) for d['printer']['name'] in list_}]

运行时出现错误:

NameError: name 'd' is not defined. Did you mean: 'id'?

这个错误是因为集合推导式语法错误,正确遍历逻辑应为for d in list_;同时仅靠name无法满足多键去重需求,还需处理idVendor和idProduct的类型差异(比如数字和字符串会被误判为不同项)。

正确的解决方法

要实现目标,需先统一三个指定键的类型,再用可哈希标识记录已出现的项,以下是两种可行实现:

方法一:遍历+集合记录

list_ = [{'printer': {'name': 'GEZHI, micro-printer', 'type': 'Usb', 'idVendor': 10473, 'idProduct': 649, 'in_ep': 129, 'out_ep': 1, 'filename': 'config1.yaml', 'status': 1}}, 
         {'printer': {'name': 'TECH, CLA58', 'type': 'Usb', 'idVendor': 26728, 'idProduct': 512, 'in_ep': 130, 'out_ep': 1, 'filename': 'config2.yaml', 'status': 0}}, 
         {'printer': {'name': 'GEZHI, micro-printer', 'type': 'Usb', 'idVendor': '10473', 'idProduct': '649', 'status': 2, 'out_ep': '1', 'in_ep': '129'}}]

seen = set()
result = []

for item in list_:
    printer_info = item['printer']
    # 统一idVendor和idProduct为字符串类型,避免类型差异导致误判
    unique_key = (
        printer_info['name'],
        str(printer_info['idVendor']),
        str(printer_info['idProduct'])
    )
    if unique_key not in seen:
        seen.add(unique_key)
        result.append(item)

print(result)

逻辑说明:

  • 用seen集合存储已出现的键组合(元组为可哈希类型,可存入集合)
  • 遍历每个元素,提取三个指定键的信息并统一类型
  • 若键组合未出现过,将当前元素加入结果列表,同时标记该组合已出现

方法二:利用字典键的唯一性

list_ = [{'printer': {'name': 'GEZHI, micro-printer', 'type': 'Usb', 'idVendor': 10473, 'idProduct': 649, 'in_ep': 129, 'out_ep': 1, 'filename': 'config1.yaml', 'status': 1}}, 
         {'printer': {'name': 'TECH, CLA58', 'type': 'Usb', 'idVendor': 26728, 'idProduct': 512, 'in_ep': 130, 'out_ep': 1, 'filename': 'config2.yaml', 'status': 0}}, 
         {'printer': {'name': 'GEZHI, micro-printer', 'type': 'Usb', 'idVendor': '10473', 'idProduct': '649', 'status': 2, 'out_ep': '1', 'in_ep': '129'}}]

unique_items = {}
for item in list_:
    printer_info = item['printer']
    unique_key = (
        printer_info['name'],
        str(printer_info['idVendor']),
        str(printer_info['idProduct'])
    )
    # 仅保留第一次出现的元素
    if unique_key not in unique_items:
        unique_items[unique_key] = item

result = list(unique_items.values())
print(result)

逻辑说明:

  • 用字典的键存储唯一标识,值对应第一个出现的元素
  • 最后将字典的值转为列表,即为去重后的结果,逻辑更简洁

内容的提问来源于stack exchange,提问作者Eduardo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 14:35:11