You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何让nltk.sent_tokenize不拆分引号内的句子?

解决NLTK分句时拆分引号内句子的问题

问题场景

使用nltk.sent_tokenize拆分句子时,引号内的多个句子会被错误拆分:

import nltk
tokenize = nltk.tokenize

t = '«I go there. Yes I go there», said John. Then he went there.'
result = tokenize.sent_tokenize(t)
print(result)
# 输出: ['«I go there.', 'Yes I go there», said John.', 'Then he went there.']

但预期结果应为:

# 期望输出: ['«I go there. Yes I go there», said John.', 'Then he went there.']

解决办法

方法1:临时替换引号内的句号规避误拆分

通过正则表达式先将引号内的句号替换为临时占位符,完成分句后再还原,避免NLTK将引号内的句号识别为分句边界:

import nltk
import re

def quote_safe_sent_tokenize(text):
    # 替换«...»引号内的句号为临时标记
    temp_text = re.sub(r'«(.*?)»', lambda match: match.group(0).replace('.', '###TEMP_DOT###'), text)
    # 执行分句
    sentences = nltk.sent_tokenize(temp_text)
    # 还原句号
    return [sent.replace('###TEMP_DOT###', '.') for sent in sentences]

# 测试
t = '«I go there. Yes I go there», said John. Then he went there.'
print(quote_safe_sent_tokenize(t))
# 输出: ['«I go there. Yes I go there», said John.', 'Then he went there.']

方法2:使用spaCy分词器

spaCy的预训练模型对上下文语义识别更精准,能自动判断引号内的句子不应拆分:

import spacy

# 加载英文预训练模型
nlp = spacy.load('en_core_web_sm')

t = '«I go there. Yes I go there», said John. Then he went there.'
doc = nlp(t)
sentences = [sent.text for sent in doc.sents]
print(sentences)
# 输出: ['«I go there. Yes I go there», said John.', 'Then he went there.']

内容的提问来源于stack exchange,提问作者Alexey Bogdanov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 23:33:29