You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何依据元组元素对嵌套元组列表重新分组并分离特定元组?

拆分嵌套元组列表中的标点元组

针对你提出的需求——把嵌套元组列表里带有'Punct'标签的元组单独提取,同时将其前后的元组分别归类到前后组中——咱们可以用两种思路实现:一种是只处理第一个出现的Punct(匹配你给出的示例),另一种是处理所有Punct元组(适用于有多个标点的场景)。


方案一:处理第一个Punct元组

这个方案完全匹配你给出的示例输出,找到第一个带有'Punct'标签的元组后,拆分前后部分并重组结构:

def split_at_first_punct(input_lol):
    pre_punct = []
    punct_tuple = None
    post_punct = []
    found_punct = False

    for sublist in input_lol:
        if not found_punct:
            # 遍历当前子列表,寻找第一个Punct
            for idx, (token, tag) in enumerate(sublist):
                if tag == 'Punct':
                    # 把Punct之前的部分加入前组
                    if idx > 0:
                        pre_punct.append(sublist[:idx])
                    # 单独记录这个Punct元组
                    punct_tuple = sublist[idx]
                    # 把Punct之后的部分加入后组
                    if idx < len(sublist) - 1:
                        post_punct.append(sublist[idx+1:])
                    found_punct = True
                    break
            else:
                # 当前子列表没有Punct,直接加入前组
                pre_punct.append(sublist)
        else:
            # 已经找到Punct,后续子列表直接加入后组
            post_punct.append(sublist)

    # 组装最终结果
    result = []
    if pre_punct:
        result.append(pre_punct)
    if punct_tuple is not None:
        result.append(punct_tuple)
    if post_punct:
        result.append(post_punct)
    
    return result

# 测试你的输入
input_lol = [ 
    [('x', 'AA'), ('y', 'AB')], 
    [('yy', 'AB'), ('..', 'Punct'), ('foo', 'ZZ')], 
    [('y', 'AB')] 
]

output = split_at_first_punct(input_lol)
print(output)

运行结果

[
    [('x', 'AA'), ('y', 'AB')], [('yy', 'AB')]
], ('..', 'Punct'), [
    [('foo', 'ZZ')], [('y', 'AB')]
]

这和你给出的示例desired_lolol结构完全一致。


方案二:处理所有Punct元组

如果你的输入里有多个'Punct'标签的元组,这个方案会逐个拆分每个标点,将连续的非标点块合并为一组:

def split_all_punct(input_lol):
    result = []
    current_group = []

    for sublist in input_lol:
        # 找出当前子列表中所有Punct的位置
        punct_indices = [i for i, (_, tag) in enumerate(sublist) if tag == 'Punct']
        start_idx = 0

        for pos in punct_indices:
            # 处理Punct之前的非标点部分
            if pos > start_idx:
                current_group.append(sublist[start_idx:pos])
                result.append(current_group)
                current_group = []
            # 单独加入Punct元组
            result.append(sublist[pos])
            # 更新起始位置,处理后续内容
            start_idx = pos + 1
        
        # 处理当前子列表剩余的非标点部分
        if start_idx < len(sublist):
            current_group.append(sublist[start_idx:])
    
    # 加入最后剩余的非标点组
    if current_group:
        result.append(current_group)
    
    return result

# 测试多个Punct的场景
test_input = [
    [('a', 'A'), ('b', 'Punct'), ('c', 'B')],
    [('d', 'C'), ('e', 'Punct'), ('f', 'D')],
    [('g', 'E')]
]

print(split_all_punct(test_input))

运行结果

[
    [('a', 'A')],
    ('b', 'Punct'),
    [('c', 'B'), ('d', 'C')],
    ('e', 'Punct'),
    [('f', 'D'), ('g', 'E')]
]

核心逻辑说明

  • 两种方案都通过遍历输入列表,识别'Punct'标签的元组位置;
  • 非标点的元组会被保留原有的子列表结构,连续的非标点块(跨原列表的)会被合并到同一个组中;
  • 每个'Punct'元组都会被单独提取,作为结果中的独立元素。

内容的提问来源于stack exchange,提问作者alvas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:40:45