关于spaCy 3.4分句工具在无标点文本上表现不佳的技术问询
spaCy无标点文本分句失效问题排查
我尝试了spaCy四种分句方案中的两种,发现它们在无标点文本上的表现同样糟糕。我希望将这类方案用于未进行说话人分轨的混合文本,目标是识别句子边界,原本认为句法分析功能能顺利将文本拆分为独立句子单元。
环境信息
python version and spacy version with language models: ============================== Info about spaCy ============================== spaCy version 3.4.3 Location /opt/homebrew/lib/python3.10/site-packages/spacy Platform macOS-12.6-arm64-arm-64bit Python version 3.10.8 Pipelines en_core_web_sm (3.4.1), en_core_web_trf (3.4.1)
我已卸载并重装spaCy及对应语言模型更新,尝试了以下两种方法:
- 依赖解析器:根据官方文档,该工具在通用新闻或网页文本上表现良好。示例代码:
nlp = spacy.load("en_core_web_sm") doc = nlp("perfect how are you doing i'm ok good to hear that can you explain me a little bit more about the situation that you send me by email") for sent in doc.sents: print(sent.text) print(token.text for token in doc)
返回结果为整段输入文本,未识别任何分句边界。
- 统计分句器(senter):根据文档,该统计模型仅提供句子边界识别功能。示例代码:
nlp = spacy.load("en_core_web_sm", exclude=["parser"]) nlp.enable_pipe("senter") doc = nlp("perfect how are you doing i'm ok good to hear that can you explain me a little bit more about the situation that you send me by email") for sent in doc.sents: print(sent.text)
返回结果与上述一致,未识别分句边界。
官方文档指出这些模型需训练后的流水线才能提供准确预测,我使用的是spaCy官方英文模型。我原以为句法分析(NP、VP等)至少能识别一个句子边界,但无标点时输入文本始终被视为单一跨度。此外,我尝试使用en_core_web_trf模型,但存在环境无法识别已安装模型的问题,此为独立问题。现咨询是否存在操作遗漏或使用错误?
内容的提问来源于stack exchange,提问作者Ja4H3ad
相关产品推荐
相关产品推荐

