Python中合并含重复URL的对象列表models字段方法问询
解决URL重复并合并Models的问题
问题描述
我有一个包含9个对象的列表,每个对象结构示例如下:
{"version": "4.6.1604", "models": ["N3NVRPOE"], "filename": "v4.6.1604.0000.206.0.1.39.3_20220316", "url": "http://files.northernvideo.com/FTP/NTH/N3NVR%20FIRMWARE/4-8-16ch/v4.6.1604.0000.206.0.1.39.3_20220316", "user_manual": "http://www.northernvideo.com/pdf/NTH-N3NVRPOE-SERIES_v1.pdf", "vendor_metadata": {"product_family": null, "model": "New N3 NVRs & Kits", "northern_status": null, "os": null, "landing_urls": ["http://www.northernvideo.com/newn3nvrskits.html", "http://northernvideo.com/newn2sdipcameras.html"]}, "description": "N3 Series H.265 -4, 8, 16 & 32 Channel NVR's with POE ", "device_picture_urls": "http://www.northernvideo.com/images/430_N3NVR_Stacked_Cropped.png"}
需求是遍历列表,删除URL重复的对象,将重复对象的models字段合并到保留对象的models列表中(例如6个同URL的对象,保留1个,其models包含所有6个对象的models)。之前尝试插入时用if验证但未成功填充主列表,寻求解决思路。
解决思路
用字典分组是最高效的方案:以URL为键,临时存储合并后的对象。遍历原列表时,根据URL是否已存在,决定是合并models还是新增对象,最后将字典的值转换为列表即可。
代码示例(Python)
# 假设原对象列表为device_list processed_dict = {} for device in device_list: current_url = device["url"] if current_url in processed_dict: # 合并models,可选去重(根据需求调整) existing_models = processed_dict[current_url]["models"] # 若不需要去重,直接用 existing_models.extend(device["models"]) merged_models = list(set(existing_models + device["models"])) processed_dict[current_url]["models"] = merged_models else: # 复制对象避免修改原数据,深拷贝用 copy.deepcopy()(若嵌套结构复杂) processed_dict[current_url] = device.copy() # 转换为最终的处理后列表 processed_list = list(processed_dict.values())
关键要点
- 字典分组的时间复杂度为O(n),远优于嵌套遍历的O(n²),适合处理任意规模的列表
- 合并models时可选择是否去重:需要去重用集合转换,不需要则直接extend
- 复制对象是为了避免修改原列表中的原始数据,防止意外副作用
- 最终将字典的值转为列表,得到去重并合并后的结果
内容的提问来源于stack exchange,提问作者Rene A
相关产品推荐
相关产品推荐

