如何去除Plotly GUI返回的字典列表中的重复项?
Plotly选中数据字典列表去重方案
问题背景
通过Plotly GUI让用户选择数据点,每次点击单个或一组数据时,对应数据会以字典形式追加到LOCAL["selected_data"]列表中,但重复点击会产生冗余字典(示例如下):
[ { "clicked": true, "selected": true, "hovered": false, "x": 0, "y": 71100.0988957607, "selected_xcol": "injection_id", "xvalue": "e54112f9-4497-4a7e-91cd-e26842a4092f", "selected_ycol": "peak_area", "yvalue": 71100.0988957607, "injection_id": "e54112f9-4497-4a7e-91cd-e26842a4092f" }, { "clicked": true, "selected": true, "hovered": false, "x": 0, "y": 75283.2386064552, "selected_xcol": "injection_id", "xvalue": "e54112f9-4497-4a7e-91cd-e26842a4092f", "selected_ycol": "peak_area", "yvalue": 75283.2386064552, "injection_id": "e54112f9-4497-4a7e-91cd-e26842a4092f" }, { # 冗余项,与第一个元素完全相同 "clicked": true, "selected": true, "hovered": false, "x": 0, "y": 71100.0988957607, "selected_xcol": "injection_id", "xvalue": "e54112f9-4497-4a7e-91cd-e26842a4092f", "selected_ycol": "peak_area", "yvalue": 71100.0988957607, "injection_id": "e54112f9-4497-4a7e-91cd-e26842a4092f" } ]
当前追加数据的代码:
LOCAL["selected_data"] += selectable_data_chart(LOCAL["df"], key = "st_react_plotly_control_main_chart", custom_data_columns = custom_data_columns, hovertemplate = hovertemplate, svgfilename = svgfilename)
失败的去重尝试
使用
set去重时触发错误:LOCAL["selected_data"] = list(set(LOCAL["selected_data"]))错误信息:
TypeError: unhashable type: 'dict'(字典不可哈希,无法存入集合)使用列表推导式结合
append时返回全null:result = [] LOCAL["selected_data"] = [result.append(d) for d in LOCAL["selected_data"] if d not in result]原因是
list.append()方法返回None,列表推导式最终收集的是一堆None值。
可行的去重方案
方案1:基于唯一标识字段去重
观察字典结构,xvalue+yvalue(或injection_id+yvalue)可作为唯一标识,遍历列表时只保留首次出现的项:
seen = set() unique_data = [] for d in LOCAL["selected_data"]: # 组合唯一标识,可根据实际业务字段调整 key = (d["xvalue"], d["yvalue"]) if key not in seen: seen.add(key) unique_data.append(d) LOCAL["selected_data"] = unique_data
方案2:基于全键值对匹配去重
如果需要严格匹配字典所有键值对来判断重复,可以把字典的键值对排序后转为可哈希的元组,再用集合去重:
# 将每个字典转为排序后的键值对元组,确保键的顺序不影响判断 unique_tuples = {tuple(sorted(d.items())) for d in LOCAL["selected_data"]} # 再转回字典列表 LOCAL["selected_data"] = [dict(t) for t in unique_tuples]
注意:此方法仅适用于字典所有值均为可哈希类型(如布尔、数字、字符串)的场景,你的示例数据符合要求。
方案3:从源头避免冗余(优化追加逻辑)
每次获取新选中的数据后,先判断是否已存在于列表中再追加,而非先追加再去重,效率更高:
new_selected = selectable_data_chart(LOCAL["df"], key = "st_react_plotly_control_main_chart", custom_data_columns = custom_data_columns, hovertemplate = hovertemplate, svgfilename = svgfilename) # 用集合存储已有的唯一标识,避免每次遍历列表判断 seen = set((d["xvalue"], d["yvalue"]) for d in LOCAL["selected_data"]) for d in new_selected: key = (d["xvalue"], d["yvalue"]) if key not in seen: seen.add(key) LOCAL["selected_data"].append(d)
内容的提问来源于stack exchange,提问作者RightmireM
相关产品推荐
相关产品推荐

