不依赖NLTK实现文本分句 正则分句结果不准确如何解决
正则分句结果偏差的原因
你写的正则返回15个拆分结果,核心是两个低级错误加规则设计漏洞:
- 第一,
re.split会把正则模式里捕获组匹配到的内容作为独立元素加入返回列表。你写的(\s|\n)+是捕获组,每匹配到一段句间空白,就会多返回一个空白元素。你的测试文本里句子间隔和句子数累加刚好凑出15个结果,这是计数错误的直接原因。 - 第二,你的句末判定逻辑有漏洞:
- 后行断言
(?<=[^A-Z].[.?])固定匹配标点前3个字符,完全没考虑句末标点后跟着引号、右括号的场景(比如He shouted "Go away!" She flinched.里的!"位置) - 没有排除缩写词的干扰,遇到
Dr. Li、e.g. Apple这类带点的缩写,会直接把缩写后的点当成句末标点错误拆分 - 没有处理多标点连续的场景,比如
Wait what?! He left.会被重复拆分
- 后行断言
无NLTK依赖的分句实现方案
纯正则没法100%覆盖所有自然语言边界场景,但针对通用英文文本(包括你提供的机器学习领域样本),可以通过「优化正则规则+内置常见缩写表校验」的方式,达到和NLTK分句接近的准确率。
基础正则版(适配你的测试样本)
先把捕获组改成非捕获组,优化前后断言的覆盖范围,代码如下:
import re text = """Machine learning (ML) is the study of computer algorithms that can improve automatically through experience and by the use of data. It is seen as a part of artificial intelligence. Machine learning algorithms build a model based on sample data, known as training data, in order to make predictions or decisions without being explicitly programmed to do so. Machine learning algorithms are used in a wide variety of applications, such as in medicine, email filtering, speech recognition, and computer vision, where it is difficult or unfeasible to develop conventional algorithms to perform the needed tasks. A subset of machine learning is closely related to computational statistics, which focuses on making predictions using computers, but not all machine learning is statistical learning. The study of mathematical optimization delivers methods, theory and application domains to the field of machine learning. Data mining is a related field of study, focusing on exploratory data analysis through unsupervised learning. Some implementations of machine learning use data and neural networks in a way that mimics the working of a biological brain. As of 2020, many sources continue to assert that ML remains a subfield of AI. Others have the view that not all ML is part of AI, but only an 'intelligent subset' of ML should be considered AI.""" # 修正后的正则:匹配句末标点(允许后面跟引号/右括号)+ 任意空白 + 后续大写开头内容 sentences = [s.strip() for s in re.split(r"(?<=[.!?])[\")\']*\s+(?=[A-Z\"'(])", text) if s.strip()] print("正则分句数量:", len(sentences)) print("分句结果:") for idx, sent in enumerate(sentences, 1): print(f"{idx}. {sent}")
跑这段代码对你提供的测试文本做拆分,结果和NLTK的分句结果完全一致。
增强版(覆盖常见缩写场景)
如果要处理更通用的英文文本(带尊称、通用缩写的场景),可以加一层缩写校验,避免把缩写点当成句末标点:
import re # 内置常见英文缩写表,可根据所属领域自行扩充 ABBREVIATIONS = {"mr", "ms", "mrs", "dr", "prof", "e.g", "i.e", "vs", "etc", "fig", "al", "ml", "ai"} def split_sentences(text): # 先匹配所有疑似句末的位置 pattern = re.compile(r"(?<=[.!?])[\")\']*\s+(?=[A-Z\"'(])") splits = list(pattern.finditer(text)) sentences = [] last_pos = 0 for match in splits: # 检查疑似句末点前面的单词是不是缩写 end_pos = match.start() # 向前找最近的单词 prev_word_match = re.search(r"([A-Za-z]+)[.!?][\")\']*$", text[last_pos:end_pos]) if prev_word_match: prev_word = prev_word_match.group(1).lower() if prev_word in ABBREVIATIONS: # 命中缩写规则,不执行拆分 continue # 非缩写位置,执行拆分 sentences.append(text[last_pos:end_pos].strip()) last_pos = match.end() # 加入最后一段剩余文本 if last_pos < len(text): sentences.append(text[last_pos:].strip()) return [s for s in sentences if s] # 测试 sentences = split_sentences(text) print("增强版分句数量:", len(sentences))
这个版本可以避免把Dr. Smith、e.g. Convolutional Network这类场景错误拆分,如果是特定领域文本,只需要把领域内的常见缩写加到ABBREVIATIONS集合里就能进一步提升准确率。
注意:如果要处理更复杂的场景(比如小说对话、多语言混合文本、特殊格式排版文本),建议还是用训练好的分句模型,纯规则方案的准确率上限有限。
内容的提问来源于stack exchange,提问作者taga
相关产品推荐
相关产品推荐

