如何解决NLTK将带引号语句错误拆分为两个分句的问题?
NLTK sent_tokenize 拆分带引号语句的问题修复
问题场景
使用NLTK的sent_tokenize处理带引号的语句时,出现了不必要的分句:
代码示例:
import nltk x = '"What do you mean?" asked Jack, looking down.' nltk.tokenize.sent_tokenize(x)
当前输出:
['"What do you mean?"', 'asked Jack, looking down.']
期望输出:
['"What do you mean?" asked Jack, looking down.']
修复方法
方法1:自定义Punkt分词器规则
NLTK的sent_tokenize基于Punkt模型,可通过添加例外规则避免引号内的标点触发分句:
import nltk from nltk.tokenize.punkt import PunktSentenceTokenizer, PunktParameters # 初始化参数,添加引号作为例外标记 punkt_param = PunktParameters() punkt_param.abbrev_types.add('"') # 创建自定义分词器并处理文本 tokenizer = PunktSentenceTokenizer(punkt_param) x = '"What do you mean?" asked Jack, looking down.' result = tokenizer.tokenize(x) print(result)
方法2:临时替换标点再还原
针对特定格式的文本,先替换引号内的分句触发标点,拆分后再还原:
import nltk x = '"What do you mean?" asked Jack, looking down.' # 替换引号内的问号为临时标记,避免触发分句 temp_text = x.replace('?"', 'TEMP_TAG"') # 执行分句 split_sentences = nltk.tokenize.sent_tokenize(temp_text) # 还原标记 fixed_sentences = [s.replace('TEMP_TAG', '?') for s in split_sentences] print(fixed_sentences)
方法3:正则合并错误拆分的句子
如果已经得到错误拆分的结果,可通过正则判断并合并相邻的引号句与后续叙述句:
import re split_sentences = ['"What do you mean?"', 'asked Jack, looking down.'] fixed = [] i = 0 while i < len(split_sentences): # 判断当前句子是否以引号结尾,且存在下一句 if re.match(r'.*"$', split_sentences[i]) and i + 1 < len(split_sentences): fixed.append(f"{split_sentences[i]} {split_sentences[i+1]}") i += 2 else: fixed.append(split_sentences[i]) i += 1 print(fixed)
内容的提问来源于stack exchange,提问作者tuvchr55
相关产品推荐
相关产品推荐

