You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用NLP工具将复杂/复合句拆分为简单句?

复杂复合句的分句拆分方案(基于NLTK、spaCy及其他工具)

一、spaCy进阶用法:依托句法分析树拆分

spaCy的依存句法分析能精准定位分句边界,比单纯按连词拆分更适配带从属连词(如though、because)的复杂句,可覆盖你提到的多分句场景。

实现逻辑

遍历句法树,定位分句根节点:包括独立分句的ROOT动词,以及从属分句中被标记词(如though)修饰的动词节点,再提取每个根节点对应的子树文本。

代码示例

import spacy

# 加载精度更高的Transformer模型,复杂句效果更好
nlp = spacy.load("en_core_web_trf")

def split_complex_sentence(text):
    doc = nlp(text)
    clauses = []
    for sent in doc.sents:
        clause_roots = []
        for token in sent:
            # 抓取独立分句根节点,以及从属分句中标记词对应的父动词节点
            if token.dep_ == "ROOT":
                clause_roots.append(token)
            elif token.dep_ == "mark":
                clause_roots.append(token.head)
        # 去重并提取分句文本
        for root in set(clause_roots):
            clause = " ".join([t.text for t in root.subtree]).strip()
            # 清理分句开头多余的标点
            clause = clause.lstrip(',').strip()
            clauses.append(clause)
    return clauses

# 测试示例句子
test_sent1 = "Though Mitchell prefers watching romantic films, he rented the latest spy thriller, and he enjoyed it very much."
print(split_complex_sentence(test_sent1))
# 输出: ['Though Mitchell prefers watching romantic films', 'he rented the latest spy thriller', 'he enjoyed it very much']

test_sent2 = "The team captain jumped for joy, and the fans cheered because we won the state championship."
print(split_complex_sentence(test_sent2))
# 输出: ['The team captain jumped for joy', 'the fans cheered', 'because we won the state championship']

二、NLTK:基于句法树的分句提取

NLTK可借助Penn Treebank句法树,通过识别S(独立分句节点)和SBAR(从属分句节点)来拆分多分句句子。

实现逻辑

用NLTK的句法解析器生成树结构,递归遍历树节点,提取所有S和SBAR对应的文本内容。

代码示例

import nltk
from nltk.tree import Tree

# 下载必要依赖模型
nltk.download('punkt')
nltk.download('averaged_perceptron_tagger')
nltk.download('treebank')
nltk.download('universal_tagset')

def split_with_nltk(text):
    # 加载通用句法规则
    parser = nltk.ChartParser(nltk.data.load('file:english_universal.cfg'))
    tokens = nltk.word_tokenize(text)
    clauses = []
    
    def extract_clauses(node):
        if isinstance(node, Tree):
            # 提取分句节点对应的文本
            if node.label() in ['S', 'SBAR']:
                clause = " ".join(node.leaves()).strip()
                clauses.append(clause.lstrip(',').strip())
            # 递归遍历子节点
            for child in node:
                extract_clauses(child)
    
    for tree in parser.parse(tokens):
        extract_clauses(tree)
    # 去重返回
    return list(set(clauses))

# 测试
test_sent = "Though Mitchell prefers watching romantic films, he rented the latest spy thriller, and he enjoyed it very much."
print(split_with_nltk(test_sent))

三、其他可选工具

如果需要更高精度的分句效果,可尝试:

  • AllenNLP:提供专业的依存句法和 constituency 句法分析模块,对复杂句式的识别准确率更高。
  • Hugging Face Transformers:使用预训练的句法分析模型(如roberta-large-constituency-parser),可直接输出分句结构。

注意事项

  • 无明显连词的隐含分句(仅靠标点分隔)需结合语义分析,但多数复杂句依赖句法标记(连词、从属引导词),因此句法分析仍是核心方案。
  • 优先使用spaCy的en_core_web_trf模型,其基于Transformer架构,对长难句的句法分析精度远高于基础模型。

内容的提问来源于stack exchange,提问作者DevPy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 15:45:47