如何在Python中对列表内相邻的指定词汇执行detokenize操作?
实现指定相邻词对的反分词合并操作
需求说明
有一个字符串列表,需要仅当“my”和“apple”按顺序相邻时,将二者通过反分词操作合并为“my apple”。给定输入列表:
words = ['this', 'is', 'my', 'apple', 'and', 'this', 'is', 'not', 'your', 'apple']
期望输出:
['this', 'is', 'my apple', 'and', 'this', 'is', 'not', 'your', 'apple']
解决方案
可以通过遍历列表并针对性检查目标词对的方式实现,结合TreebankWordDetokenizer完成规范的反分词合并:
完整代码实现
from nltk.tokenize.treebank import TreebankWordDetokenizer # 初始化反分词器 detokenizer = TreebankWordDetokenizer() def merge_specific_pair(word_list, target_pair): result = [] index = 0 total_words = len(word_list) while index < total_words: # 检查当前位置和下一个位置是否匹配目标词对 if index < total_words - 1 and word_list[index] == target_pair[0] and word_list[index+1] == target_pair[1]: # 对目标词对执行反分词合并 merged_str = detokenizer.detokenize([word_list[index], word_list[index+1]]) result.append(merged_str) index += 2 # 跳过已合并的下一个元素 else: result.append(word_list[index]) index += 1 return result # 测试示例 words = ['this', 'is', 'my', 'apple', 'and', 'this', 'is', 'not', 'your', 'apple'] target = ('my', 'apple') output = merge_specific_pair(words, target) print(output)
代码逻辑说明
- 初始化
TreebankWordDetokenizer用于规范的反分词操作; - 定义处理函数,通过
while循环遍历原列表:- 每次检查当前元素和下一个元素是否匹配指定的目标词对;
- 若匹配则合并两个词并加入结果列表,同时跳过下一个元素;
- 若不匹配则直接将当前元素加入结果列表,继续遍历下一个元素;
- 最终返回处理后的结果列表。
内容的提问来源于stack exchange,提问作者Quinten
相关产品推荐
相关产品推荐

