如何使用NLTK基于句号和分号实现句子分词?
如何用NLTK的sent_tokenize基于句号和分号分句?
默认的nltk.sent_tokenize()用的是预训练的英文分句模型,只会把句号、感叹号这类标准句子结束符作为分句依据,不会识别分号。要实现按句号和分号分句,你需要自定义一个Punkt分句器,通过训练让它把分号也当作分句分隔符。
具体实现步骤:
- 导入NLTK的Punkt相关模块
- 用包含目标标点的文本训练自定义分句模型
- 使用训练好的分句器处理文本
完整代码示例:
import nltk from nltk.tokenize.punkt import PunktSentenceTokenizer, PunktTrainer # 准备训练文本,包含我们需要的分句标点(分号和句号) train_text = "The court ruled out the judgement; First round proceedings; The court declared unjustfied. Case proceedings were carried out in the morning. Objections were raised;" # 初始化训练器,开启所有搭配识别 trainer = PunktTrainer() trainer.INCLUDE_ALL_COLLOCS = True trainer.train(train_text) # 创建自定义分句器 custom_tokenizer = PunktSentenceTokenizer(trainer.get_params()) # 处理目标文本 sentence = "The court ruled out the judgement; First round proceedings; The court declared unjustfied. Case proceedings were carried out in the morning. Objections were raised;" sentences = custom_tokenizer.tokenize(sentence) print(sentences)
运行结果:
[ "The court ruled out the judgement;", "First round proceedings;", "The court declared unjustfied.", "Case proceedings were carried out in the morning.", "Objections were raised;" ]
补充说明:
如果你的语料中有更多不同场景的文本,可以把更多包含分号和句号的句子加入训练文本,这样自定义分句器的泛化性会更好。核心是通过训练让Punkt模型学习到“分号也是句子结束的标志”,从而实现你需要的分句逻辑。
内容的提问来源于stack exchange,提问作者Ashreen Kaur
相关产品推荐
相关产品推荐

