Python中按incidentId比较字典列表,返回不匹配项的方法
你的问题出在直接比较整个字典对象是否存在于另一个列表里——虽然两个列表里的条目incidentId相同,但incidentState不一样,所以每个字典都会被判定为“不在另一个列表中”,自然就把所有项都加入non_match了。
要实现基于incidentId的匹配判断,核心思路是先提取出其中一个列表里所有的incidentId(用集合存储效率最高),再逐个检查另一个列表里的条目是否包含不在这个集合里的incidentId。
方法一:找出incidents中incidentId不在incidents_db里的项
incidents = [ { "incidentId": "RWS03_949563", "incidentState": "REGISTERED" }, { "incidentId": "RWS03_949565", "incidentState": "REGISTERED" } ] incidents_db = [ { "incidentId": "RWS03_949563", "incidentState": "PENDING" }, { "incidentId": "RWS03_949565", "incidentState": "PENDING" } ] # 先把incidents_db里的所有incidentId提取到集合中,集合查询效率远高于列表 db_incident_ids = {item["incidentId"] for item in incidents_db} non_match = [] for incident in incidents: # 只比较incidentId字段 if incident["incidentId"] not in db_incident_ids: non_match.append(incident) print(non_match) # 你的例子中两个incidentId都存在,所以输出[]
方法二:双向找出两边都不匹配的项(如果需要)
如果你需要同时找出incidents里不在incidents_db的项,以及incidents_db里不在incidents的项,可以这样写:
# 提取两边的incidentId集合 incident_ids = {item["incidentId"] for item in incidents} db_incident_ids = {item["incidentId"] for item in incidents_db} # 分别筛选不匹配项 non_match_incidents = [item for item in incidents if item["incidentId"] not in db_incident_ids] non_match_db = [item for item in incidents_db if item["incidentId"] not in incident_ids] # 合并结果 all_non_matches = non_match_incidents + non_match_db
为什么用集合?
用集合存储incidentId的原因是查询时间复杂度为O(1),如果直接用列表的in操作,每次查询都是O(n),当你的列表条目很多时,集合的效率会高很多。
内容的提问来源于stack exchange,提问作者Lucas Scheepers
相关产品推荐
相关产品推荐

