You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决NLTK将带引号语句错误拆分为两个分句的问题?

NLTK sent_tokenize 拆分带引号语句的问题修复

问题场景

使用NLTK的sent_tokenize处理带引号的语句时,出现了不必要的分句:

代码示例:

import nltk
x = '"What do you mean?" asked Jack, looking down.'
nltk.tokenize.sent_tokenize(x)

当前输出:

['"What do you mean?"', 'asked Jack, looking down.']

期望输出:

['"What do you mean?" asked Jack, looking down.']

修复方法

方法1:自定义Punkt分词器规则

NLTK的sent_tokenize基于Punkt模型,可通过添加例外规则避免引号内的标点触发分句:

import nltk
from nltk.tokenize.punkt import PunktSentenceTokenizer, PunktParameters

# 初始化参数,添加引号作为例外标记
punkt_param = PunktParameters()
punkt_param.abbrev_types.add('"')

# 创建自定义分词器并处理文本
tokenizer = PunktSentenceTokenizer(punkt_param)
x = '"What do you mean?" asked Jack, looking down.'
result = tokenizer.tokenize(x)
print(result)

方法2:临时替换标点再还原

针对特定格式的文本,先替换引号内的分句触发标点,拆分后再还原:

import nltk

x = '"What do you mean?" asked Jack, looking down.'
# 替换引号内的问号为临时标记,避免触发分句
temp_text = x.replace('?"', 'TEMP_TAG"')
# 执行分句
split_sentences = nltk.tokenize.sent_tokenize(temp_text)
# 还原标记
fixed_sentences = [s.replace('TEMP_TAG', '?') for s in split_sentences]
print(fixed_sentences)

方法3:正则合并错误拆分的句子

如果已经得到错误拆分的结果,可通过正则判断并合并相邻的引号句与后续叙述句:

import re

split_sentences = ['"What do you mean?"', 'asked Jack, looking down.']
fixed = []
i = 0
while i < len(split_sentences):
    # 判断当前句子是否以引号结尾,且存在下一句
    if re.match(r'.*&quot;$', split_sentences[i]) and i + 1 < len(split_sentences):
        fixed.append(f"{split_sentences[i]} {split_sentences[i+1]}")
        i += 2
    else:
        fixed.append(split_sentences[i])
        i += 1
print(fixed)

内容的提问来源于stack exchange,提问作者tuvchr55

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 18:18:49