You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中对列表内相邻的指定词汇执行detokenize操作?

实现指定相邻词对的反分词合并操作

需求说明

有一个字符串列表,需要仅当“my”和“apple”按顺序相邻时,将二者通过反分词操作合并为“my apple”。给定输入列表:

words = ['this', 'is', 'my', 'apple', 'and', 'this', 'is', 'not', 'your', 'apple']

期望输出:

['this', 'is', 'my apple', 'and', 'this', 'is', 'not', 'your', 'apple']

解决方案

可以通过遍历列表并针对性检查目标词对的方式实现,结合TreebankWordDetokenizer完成规范的反分词合并:

完整代码实现

from nltk.tokenize.treebank import TreebankWordDetokenizer

# 初始化反分词器
detokenizer = TreebankWordDetokenizer()

def merge_specific_pair(word_list, target_pair):
    result = []
    index = 0
    total_words = len(word_list)
    
    while index < total_words:
        # 检查当前位置和下一个位置是否匹配目标词对
        if index < total_words - 1 and word_list[index] == target_pair[0] and word_list[index+1] == target_pair[1]:
            # 对目标词对执行反分词合并
            merged_str = detokenizer.detokenize([word_list[index], word_list[index+1]])
            result.append(merged_str)
            index += 2  # 跳过已合并的下一个元素
        else:
            result.append(word_list[index])
            index += 1
    return result

# 测试示例
words = ['this', 'is', 'my', 'apple', 'and', 'this', 'is', 'not', 'your', 'apple']
target = ('my', 'apple')
output = merge_specific_pair(words, target)
print(output)

代码逻辑说明

  1. 初始化TreebankWordDetokenizer用于规范的反分词操作;
  2. 定义处理函数,通过while循环遍历原列表:
    • 每次检查当前元素和下一个元素是否匹配指定的目标词对;
    • 若匹配则合并两个词并加入结果列表,同时跳过下一个元素;
    • 若不匹配则直接将当前元素加入结果列表,继续遍历下一个元素;
  3. 最终返回处理后的结果列表。

内容的提问来源于stack exchange,提问作者Quinten

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 06:52:51