You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用NLTK基于句号和分号实现句子分词?

如何用NLTK的sent_tokenize基于句号和分号分句?

默认的nltk.sent_tokenize()用的是预训练的英文分句模型,只会把句号、感叹号这类标准句子结束符作为分句依据,不会识别分号。要实现按句号和分号分句,你需要自定义一个Punkt分句器,通过训练让它把分号也当作分句分隔符。

具体实现步骤:

  • 导入NLTK的Punkt相关模块
  • 用包含目标标点的文本训练自定义分句模型
  • 使用训练好的分句器处理文本

完整代码示例:

import nltk
from nltk.tokenize.punkt import PunktSentenceTokenizer, PunktTrainer

# 准备训练文本,包含我们需要的分句标点(分号和句号)
train_text = "The court ruled out the judgement; First round proceedings; The court declared unjustfied. Case proceedings were carried out in the morning. Objections were raised;"

# 初始化训练器,开启所有搭配识别
trainer = PunktTrainer()
trainer.INCLUDE_ALL_COLLOCS = True
trainer.train(train_text)

# 创建自定义分句器
custom_tokenizer = PunktSentenceTokenizer(trainer.get_params())

# 处理目标文本
sentence = "The court ruled out the judgement; First round proceedings; The court declared unjustfied. Case proceedings were carried out in the morning. Objections were raised;"
sentences = custom_tokenizer.tokenize(sentence)

print(sentences)

运行结果:

[
 "The court ruled out the judgement;",
 "First round proceedings;",
 "The court declared unjustfied.",
 "Case proceedings were carried out in the morning.",
 "Objections were raised;"
]

补充说明:

如果你的语料中有更多不同场景的文本,可以把更多包含分号和句号的句子加入训练文本,这样自定义分句器的泛化性会更好。核心是通过训练让Punkt模型学习到“分号也是句子结束的标志”,从而实现你需要的分句逻辑。

内容的提问来源于stack exchange,提问作者Ashreen Kaur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 22:49:54