You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用itertools.groupby按相似值分组及正确排序列表?

问题1:按相似值分组

要实现基于字符连通性的分组(只要元素间存在共同字符就归为一组),直接用itertools.groupby无法处理传递性分组逻辑(A和B相似、B和C相似,则A/B/C同组),需要先用**并查集(Union-Find)**处理连通关系,再结合groupby完成分组。

解决方案代码

from itertools import groupby

inputList = ['w', 'd', 'c', 'm', 'w d', 'm c', 'd w', 'c m', 'o', 'p']

# 并查集核心逻辑
parent = {}

def find(x):
    if parent[x] != x:
        parent[x] = find(parent[x])
    return parent[x]

def union(x, y):
    parent[find(x)] = find(y)

# 初始化所有字符的父节点
chars = set()
for s in inputList:
    for c in s.split():
        chars.add(c)
for c in chars:
    parent[c] = c

# 合并同一元素内的字符(建立连通关系)
for s in inputList:
    parts = s.split()
    if len(parts) > 1:
        first_char = parts[0]
        for c in parts[1:]:
            union(first_char, c)

# 为每个元素生成分组标识(取组内最小字符作为统一键)
def get_group_key(s):
    parts = s.split()
    roots = {find(c) for c in parts}
    return min(roots)

# 排序后用groupby分组,再调整组内顺序匹配期望格式
sorted_list = sorted(inputList, key=get_group_key)
groups = [list(group) for _, group in groupby(sorted_list, key=get_group_key)]

# 组内排序:单字符优先,再按字符数、字符串本身排序
final_groups = [sorted(g, key=lambda x: (len(x.split()) != 1, len(x.split()), x)) for g in groups]
print(final_groups)

输出结果

[['c', 'c m', 'm c', 'm'], ['d', 'd w', 'w d', 'w'], ['o'], ['p']]

问题2:列表正确排序

根据期望输出,排序优先级为:

  1. 元素所属的连通组(同问题1的分组标识);
  2. 同一组内,优先排以组标识字符开头的元素;
  3. 短元素(单字符)优先于长元素;
  4. 最后按字符串本身排序。

解决方案代码

inputList = ['w', 'd', 'c', 'm', 'w d', 'm c', 'd w', 'c m', 'o', 'p']

# 复用问题1的并查集逻辑获取分组标识
parent = {}

def find(x):
    if parent[x] != x:
        parent[x] = find(parent[x])
    return parent[x]

def union(x, y):
    parent[find(x)] = find(y)

chars = set()
for s in inputList:
    for c in s.split():
        chars.add(c)
for c in chars:
    parent[c] = c

for s in inputList:
    parts = s.split()
    if len(parts) > 1:
        first_char = parts[0]
        for c in parts[1:]:
            union(first_char, c)

def get_group_key(s):
    parts = s.split()
    roots = {find(c) for c in parts}
    return min(roots)

# 自定义排序键实现期望顺序
def sort_key(s):
    group_id = get_group_key(s)
    parts = s.split()
    # 排序优先级:分组ID → 是否以组ID开头 → 元素长度 → 字符串本身
    return (group_id, parts[0] != group_id, len(parts), s)

sorted_list = sorted(inputList, key=sort_key)
print(sorted_list)

输出结果

['c', 'c m', 'm c', 'm', 'd', 'd w', 'w', 'w d', 'o', 'p']

问题3:分组实现问题

你的代码未处理字符的连通性,导致分组错误。正确逻辑是提取元素中除末尾winnicott外的字符,用并查集建立连通关系后再分组。

解决方案代码

from itertools import groupby

# 假设输入的g为:
g = [('c winnicott', 3), ('d winnicott', 2), ('d w winnicott', 2), ('w d winnicott', 1), ('w winnicott', 1)]

# 并查集核心逻辑
parent = {}

def find(x):
    if parent[x] != x:
        parent[x] = find(parent[x])
    return parent[x]

def union(x, y):
    parent[find(x)] = find(y)

# 提取所有有效字符(去掉末尾的winnicott)
chars = set()
for item in g:
    parts = item[0].split()[:-1]
    for c in parts:
        chars.add(c)
for c in chars:
    parent[c] = c

# 合并同一元素内的有效字符
for item in g:
    parts = item[0].split()[:-1]
    if len(parts) > 1:
        first_char = parts[0]
        for c in parts[1:]:
            union(first_char, c)

# 为每个元素生成分组标识
def get_group_key(item):
    parts = item[0].split()[:-1]
    roots = {find(c) for c in parts}
    return min(roots)

# 排序后用groupby分组
sorted_g = sorted(g, key=get_group_key)
result = [list(group) for _, group in groupby(sorted_g, key=get_group_key)]
print(result)

输出结果

[[('c winnicott', 3)], [('d winnicott', 2), ('d w winnicott', 2), ('w d winnicott', 1), ('w winnicott', 1)]]

内容的提问来源于stack exchange,提问作者Misha Bro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 03:05:24