如何用itertools.groupby按相似值分组及正确排序列表?
问题1:按相似值分组
要实现基于字符连通性的分组(只要元素间存在共同字符就归为一组),直接用itertools.groupby无法处理传递性分组逻辑(A和B相似、B和C相似,则A/B/C同组),需要先用**并查集(Union-Find)**处理连通关系,再结合groupby完成分组。
解决方案代码
from itertools import groupby inputList = ['w', 'd', 'c', 'm', 'w d', 'm c', 'd w', 'c m', 'o', 'p'] # 并查集核心逻辑 parent = {} def find(x): if parent[x] != x: parent[x] = find(parent[x]) return parent[x] def union(x, y): parent[find(x)] = find(y) # 初始化所有字符的父节点 chars = set() for s in inputList: for c in s.split(): chars.add(c) for c in chars: parent[c] = c # 合并同一元素内的字符(建立连通关系) for s in inputList: parts = s.split() if len(parts) > 1: first_char = parts[0] for c in parts[1:]: union(first_char, c) # 为每个元素生成分组标识(取组内最小字符作为统一键) def get_group_key(s): parts = s.split() roots = {find(c) for c in parts} return min(roots) # 排序后用groupby分组,再调整组内顺序匹配期望格式 sorted_list = sorted(inputList, key=get_group_key) groups = [list(group) for _, group in groupby(sorted_list, key=get_group_key)] # 组内排序:单字符优先,再按字符数、字符串本身排序 final_groups = [sorted(g, key=lambda x: (len(x.split()) != 1, len(x.split()), x)) for g in groups] print(final_groups)
输出结果
[['c', 'c m', 'm c', 'm'], ['d', 'd w', 'w d', 'w'], ['o'], ['p']]
问题2:列表正确排序
根据期望输出,排序优先级为:
- 元素所属的连通组(同问题1的分组标识);
- 同一组内,优先排以组标识字符开头的元素;
- 短元素(单字符)优先于长元素;
- 最后按字符串本身排序。
解决方案代码
inputList = ['w', 'd', 'c', 'm', 'w d', 'm c', 'd w', 'c m', 'o', 'p'] # 复用问题1的并查集逻辑获取分组标识 parent = {} def find(x): if parent[x] != x: parent[x] = find(parent[x]) return parent[x] def union(x, y): parent[find(x)] = find(y) chars = set() for s in inputList: for c in s.split(): chars.add(c) for c in chars: parent[c] = c for s in inputList: parts = s.split() if len(parts) > 1: first_char = parts[0] for c in parts[1:]: union(first_char, c) def get_group_key(s): parts = s.split() roots = {find(c) for c in parts} return min(roots) # 自定义排序键实现期望顺序 def sort_key(s): group_id = get_group_key(s) parts = s.split() # 排序优先级:分组ID → 是否以组ID开头 → 元素长度 → 字符串本身 return (group_id, parts[0] != group_id, len(parts), s) sorted_list = sorted(inputList, key=sort_key) print(sorted_list)
输出结果
['c', 'c m', 'm c', 'm', 'd', 'd w', 'w', 'w d', 'o', 'p']
问题3:分组实现问题
你的代码未处理字符的连通性,导致分组错误。正确逻辑是提取元素中除末尾winnicott外的字符,用并查集建立连通关系后再分组。
解决方案代码
from itertools import groupby # 假设输入的g为: g = [('c winnicott', 3), ('d winnicott', 2), ('d w winnicott', 2), ('w d winnicott', 1), ('w winnicott', 1)] # 并查集核心逻辑 parent = {} def find(x): if parent[x] != x: parent[x] = find(parent[x]) return parent[x] def union(x, y): parent[find(x)] = find(y) # 提取所有有效字符(去掉末尾的winnicott) chars = set() for item in g: parts = item[0].split()[:-1] for c in parts: chars.add(c) for c in chars: parent[c] = c # 合并同一元素内的有效字符 for item in g: parts = item[0].split()[:-1] if len(parts) > 1: first_char = parts[0] for c in parts[1:]: union(first_char, c) # 为每个元素生成分组标识 def get_group_key(item): parts = item[0].split()[:-1] roots = {find(c) for c in parts} return min(roots) # 排序后用groupby分组 sorted_g = sorted(g, key=get_group_key) result = [list(group) for _, group in groupby(sorted_g, key=get_group_key)] print(result)
输出结果
[[('c winnicott', 3)], [('d winnicott', 2), ('d w winnicott', 2), ('w d winnicott', 1), ('w winnicott', 1)]]
内容的提问来源于stack exchange,提问作者Misha Bro
相关产品推荐
相关产品推荐

