如何让nltk.sent_tokenize不拆分引号内的句子?
解决NLTK分句时拆分引号内句子的问题
问题场景
使用nltk.sent_tokenize拆分句子时,引号内的多个句子会被错误拆分:
import nltk tokenize = nltk.tokenize t = '«I go there. Yes I go there», said John. Then he went there.' result = tokenize.sent_tokenize(t) print(result) # 输出: ['«I go there.', 'Yes I go there», said John.', 'Then he went there.']
但预期结果应为:
# 期望输出: ['«I go there. Yes I go there», said John.', 'Then he went there.']
解决办法
方法1:临时替换引号内的句号规避误拆分
通过正则表达式先将引号内的句号替换为临时占位符,完成分句后再还原,避免NLTK将引号内的句号识别为分句边界:
import nltk import re def quote_safe_sent_tokenize(text): # 替换«...»引号内的句号为临时标记 temp_text = re.sub(r'«(.*?)»', lambda match: match.group(0).replace('.', '###TEMP_DOT###'), text) # 执行分句 sentences = nltk.sent_tokenize(temp_text) # 还原句号 return [sent.replace('###TEMP_DOT###', '.') for sent in sentences] # 测试 t = '«I go there. Yes I go there», said John. Then he went there.' print(quote_safe_sent_tokenize(t)) # 输出: ['«I go there. Yes I go there», said John.', 'Then he went there.']
方法2:使用spaCy分词器
spaCy的预训练模型对上下文语义识别更精准,能自动判断引号内的句子不应拆分:
import spacy # 加载英文预训练模型 nlp = spacy.load('en_core_web_sm') t = '«I go there. Yes I go there», said John. Then he went there.' doc = nlp(t) sentences = [sent.text for sent in doc.sents] print(sentences) # 输出: ['«I go there. Yes I go there», said John.', 'Then he went there.']
内容的提问来源于stack exchange,提问作者Alexey Bogdanov
相关产品推荐
相关产品推荐

