You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Swift使用NaturalLanguage分割学术文本句子的错误处理方法

问题描述

我当前使用Swift 5,需要实现学术文本的独立句子拆分功能。
NaturalLanguage框架对绝大多数通用文本的句子分割效果优异,但无法正确处理学术文本中的特殊内容,典型错误触发场景包括et al.缩写、页码标记"p. "、连续括号()()模式,会在这些位置错误切分句子。

可复现代码
import NaturalLanguage

var sentences: [String] = []
var str = "The information was reported by Brown et al. (2000). This should not have been the case (Brown et al., 2001, p. 10). Several other studies corroborate this (i.e., I don't know but something important, etc.) (Brown et al., 2002; Brown et al., 2003). But this is weird given the results of White et al. (2001)."

str.enumerateSubstrings(in: str.startIndex..., options: [.localized, .bySentences]) { (tag, _, _, _) in
           sentences.append(tag ?? "")
       }

sentences.forEach {
    print($0)
}


print(sentences)
预期分割结果
The information was reported by Brown et al. (2000). 
This should not have been the case (Brown et al., 2001, p. 10). 
Several other studies corroborate this (i.e., I don't know but something important, etc.) (Brown et al., 2002; Brown et al., 2003). 
But this is weird given the results of White et al. (2001).
实际运行结果
The information was reported by Brown et al. 
(2000). 
This should not have been the case (Brown et al., 2001, p. 
10). 
Several other studies corroborate this (i.e., I don't know but something important, etc.) 
(Brown et al., 2002; Brown et al., 2003). 
But this is weird given the results of White et al. 
(2001).


["The information was reported by Brown et al. ", "(2000). ", "This should not have been the case (Brown et al., 2001, p. ", "10). ", "Several other studies corroborate this (i.e., I don\'t know but something important, etc.) ", "(Brown et al., 2002; Brown et al., 2003). ", "But this is weird given the results of White et al. ", "(2001)."]
解决方法

NaturalLanguage没有开放自定义分句规则的公开接口,针对学术文本的分句错误,不需要更换基础框架,通过「预处理+后处理合并」的方案即可修复:

  • 预处理阶段:调用系统分句接口前,先把文本中所有已知会触发错误切分的学术固定写法做临时占位替换,比如把et al.替换为不带句点的临时字符串,把页码标记p. 里的句点替换为特殊无意义标记,避免分词器把这类缩写里的句点误判为句末停顿。
  • 分句阶段:正常调用NaturalLanguage的分句接口,拿到初始切分结果。
  • 后处理阶段:遍历初始切分的片段做合并校验:如果当前片段以et al.、p.这类明确不属于句末的缩写结尾,就把下一个片段拼接上来,直到碰到真正的句末标识(句末点号后接句首大写的新内容,或是完整闭合的括号后接句末点号);如果片段以未配对的左括号开头,说明是被错误切开的引用内容,直接拼接到上一个片段末尾。所有合并操作完成后,把之前替换的占位符还原为原始字符即可。

如果处理的文本规模大、特殊场景多,不想自行维护规则列表,也可以选择专门适配学术出版场景的分句库,这类库内置了全量学术缩写表、括号/引号配对校验逻辑,对APA等常见引用格式的适配效果远好于系统通用分句器,不需要手动编写大量特殊场景判断逻辑。


内容的提问来源于stack exchange,提问作者user8460166

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 17:06:28