You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取单词列表形式文本中多词目标序列的对应索引

多词目标在单词列表中匹配索引的实现方法

实现思路

  • 先把所有待匹配的多词目标拆分成单个单词的列表,同时记录每个目标的单词长度
  • 遍历上下文单词列表的每一个可能的起始位置,只要起始位置加上目标单词长度不超过上下文总长度,就取出对应长度的上下文片段和目标拆分后的列表做对比
  • 匹配成功就把当前的起始索引和结束索引(起始索引 + 目标长度 - 1)加入结果集
  • 最后可根据需要对结果按索引先后排序,避免因为目标输入顺序导致结果顺序混乱

代码实现(Python为例)

def find_multi_word_indices(context: list[str], targets: list[str]) -> list[list[int]]:
    result = []
    # 预处理目标:把每个多词目标拆成单词列表
    processed_targets = [t.split() for t in targets]
    context_len = len(context)

    for target_words in processed_targets:
        target_len = len(target_words)
        # 目标长度大于上下文总长度时直接跳过,避免无效遍历
        if target_len > context_len:
            continue
        # 遍历所有可能的起始索引
        for start_idx in range(context_len - target_len + 1):
            # 切片匹配目标单词
            if context[start_idx:start_idx+target_len] == target_words:
                end_idx = start_idx + target_len - 1
                result.append([start_idx, end_idx])
    # 按索引从小到大排序,保证结果顺序和出现顺序一致
    result.sort()
    return result

# 测试示例
context = ['Katie', 'Joplin', 'is', 'an', 'American', 'sitcom', 'created', 'by', 'Tom', 'Seeley', 'and', 'Norm', 'Gunzenhauser', '.', 'The', 'sitcom', 'received', 'positive', 'reviews', 'thanks', 'to', 'the', 'brilliance', 'of', 'Tom', 'Seeley', '.']
target = ['sitcom created', 'Tom Seeley']
print(find_multi_word_indices(context, target))

运行后输出结果为[[5, 6], [8, 9], [24, 25]],和要求的输出完全一致。

注意事项

  • 代码支持任意长度的多词目标匹配,不需要针对不同长度的短语做额外修改
  • 如果需要按目标分组返回匹配结果,可调整返回结构,用字典存储每个目标对应的索引列表
  • 不需要按出现顺序排序的话可以删除排序逻辑,结果会按目标的输入顺序返回

内容的提问来源于stack exchange,提问作者Balázs Fehér

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 16:45:04