如何获取单词列表形式文本中多词目标序列的对应索引
多词目标在单词列表中匹配索引的实现方法
实现思路
- 先把所有待匹配的多词目标拆分成单个单词的列表,同时记录每个目标的单词长度
- 遍历上下文单词列表的每一个可能的起始位置,只要起始位置加上目标单词长度不超过上下文总长度,就取出对应长度的上下文片段和目标拆分后的列表做对比
- 匹配成功就把当前的起始索引和结束索引(起始索引 + 目标长度 - 1)加入结果集
- 最后可根据需要对结果按索引先后排序,避免因为目标输入顺序导致结果顺序混乱
代码实现(Python为例)
def find_multi_word_indices(context: list[str], targets: list[str]) -> list[list[int]]: result = [] # 预处理目标:把每个多词目标拆成单词列表 processed_targets = [t.split() for t in targets] context_len = len(context) for target_words in processed_targets: target_len = len(target_words) # 目标长度大于上下文总长度时直接跳过,避免无效遍历 if target_len > context_len: continue # 遍历所有可能的起始索引 for start_idx in range(context_len - target_len + 1): # 切片匹配目标单词 if context[start_idx:start_idx+target_len] == target_words: end_idx = start_idx + target_len - 1 result.append([start_idx, end_idx]) # 按索引从小到大排序,保证结果顺序和出现顺序一致 result.sort() return result # 测试示例 context = ['Katie', 'Joplin', 'is', 'an', 'American', 'sitcom', 'created', 'by', 'Tom', 'Seeley', 'and', 'Norm', 'Gunzenhauser', '.', 'The', 'sitcom', 'received', 'positive', 'reviews', 'thanks', 'to', 'the', 'brilliance', 'of', 'Tom', 'Seeley', '.'] target = ['sitcom created', 'Tom Seeley'] print(find_multi_word_indices(context, target))
运行后输出结果为[[5, 6], [8, 9], [24, 25]],和要求的输出完全一致。
注意事项
- 代码支持任意长度的多词目标匹配,不需要针对不同长度的短语做额外修改
- 如果需要按目标分组返回匹配结果,可调整返回结构,用字典存储每个目标对应的索引列表
- 不需要按出现顺序排序的话可以删除排序逻辑,结果会按目标的输入顺序返回
内容的提问来源于stack exchange,提问作者Balázs Fehér
相关产品推荐
相关产品推荐

