如何对多维列表按匹配字符串元素排序并去除重复匹配项
解决多维列表的重复匹配去重及按首次出现值排序问题
我来帮你搞定这两个需求,咱们拆解成两个部分一步步解决:
一、获取无重复的匹配位置列表(目标1)
先看你当前代码的问题:
- 循环逻辑会产生多余的组内匹配(比如
[2,5],但其实0已经和2、5配对过了,这个配对是不必要的) - 判断条件用了
data[i][2] in data[j+1][2],这会匹配部分包含的情况,而你需要的是完全相等的字符串匹配
正确的思路是先按data[x][2]的字符串值分组,收集所有对应索引,然后只保留每组第一个索引和其他索引的配对,这样就能避免重复匹配。
实现代码:
data = [["something1", 1, "number one", "inf 1",1, 33,22, "other"], ["something2",2, "number twenty", "inf 2", 1,66, 11, "other"], ["something3",3, "number one", "inf 3", 1,99, 55, "other"], ["something4",4, "number five", "inf 4", 1, 1212, 9988, "other"], ["something5",3, "number four", "inf 3", 1,99, 55, "other"], ["something6",3, "number one", "inf 3", 1,99, 55, "other"], ["something7",3, "number twenty", "inf 3", 1,99, 55, "other"]] from collections import defaultdict # 1. 按目标字符串分组,收集所有对应索引 index_groups = defaultdict(list) for idx, item in enumerate(data): target_str = item[2] index_groups[target_str].append(idx) # 2. 生成去重的匹配列表:仅保留每组第一个索引与其他索引的配对 listMatch = [] for indices in index_groups.values(): if len(indices) >= 2: first_idx = indices[0] for other_idx in indices[1:]: listMatch.append([first_idx, other_idx]) print(listMatch) # 输出: [[0, 2], [0, 5], [1, 6]]
二、按首次出现位置排序多维列表(目标2)
要让同字符串的元素聚集在一起,且以该字符串首次出现的位置作为排序依据,我们可以先记录每个字符串的首次出现索引,再用这个索引作为排序key。
实现代码:
# 1. 记录每个目标字符串的首次出现索引 first_occurrence = {} for idx, item in enumerate(data): target_str = item[2] if target_str not in first_occurrence: first_occurrence[target_str] = idx # 2. 根据首次出现索引排序 sorted_data = sorted(data, key=lambda x: first_occurrence[x[2]]) # 打印排序后的结果(和你预期的目标2一致) print(sorted_data)
完整整合脚本
如果需要把两个功能放在一起,直接合并代码即可:
data = [["something1", 1, "number one", "inf 1",1, 33,22, "other"], ["something2",2, "number twenty", "inf 2", 1,66, 11, "other"], ["something3",3, "number one", "inf 3", 1,99, 55, "other"], ["something4",4, "number five", "inf 4", 1, 1212, 9988, "other"], ["something5",3, "number four", "inf 3", 1,99, 55, "other"], ["something6",3, "number one", "inf 3", 1,99, 55, "other"], ["something7",3, "number twenty", "inf 3", 1,99, 55, "other"]] from collections import defaultdict # 目标1:生成去重的匹配位置列表 index_groups = defaultdict(list) for idx, item in enumerate(data): target_str = item[2] index_groups[target_str].append(idx) listMatch = [] for indices in index_groups.values(): if len(indices) >= 2: first_idx = indices[0] for other_idx in indices[1:]: listMatch.append([first_idx, other_idx]) print("去重后的匹配列表:", listMatch) # 目标2:按首次出现位置排序 first_occurrence = {} for idx, item in enumerate(data): target_str = item[2] if target_str not in first_occurrence: first_occurrence[target_str] = idx sorted_data = sorted(data, key=lambda x: first_occurrence[x[2]]) print("排序后的多维列表:", sorted_data)
内容的提问来源于stack exchange,提问作者Pr.Syn
相关产品推荐
相关产品推荐

