如何用NLP工具将复杂/复合句拆分为简单句?
复杂复合句的分句拆分方案(基于NLTK、spaCy及其他工具)
一、spaCy进阶用法:依托句法分析树拆分
spaCy的依存句法分析能精准定位分句边界,比单纯按连词拆分更适配带从属连词(如though、because)的复杂句,可覆盖你提到的多分句场景。
实现逻辑
遍历句法树,定位分句根节点:包括独立分句的ROOT动词,以及从属分句中被标记词(如though)修饰的动词节点,再提取每个根节点对应的子树文本。
代码示例
import spacy # 加载精度更高的Transformer模型,复杂句效果更好 nlp = spacy.load("en_core_web_trf") def split_complex_sentence(text): doc = nlp(text) clauses = [] for sent in doc.sents: clause_roots = [] for token in sent: # 抓取独立分句根节点,以及从属分句中标记词对应的父动词节点 if token.dep_ == "ROOT": clause_roots.append(token) elif token.dep_ == "mark": clause_roots.append(token.head) # 去重并提取分句文本 for root in set(clause_roots): clause = " ".join([t.text for t in root.subtree]).strip() # 清理分句开头多余的标点 clause = clause.lstrip(',').strip() clauses.append(clause) return clauses # 测试示例句子 test_sent1 = "Though Mitchell prefers watching romantic films, he rented the latest spy thriller, and he enjoyed it very much." print(split_complex_sentence(test_sent1)) # 输出: ['Though Mitchell prefers watching romantic films', 'he rented the latest spy thriller', 'he enjoyed it very much'] test_sent2 = "The team captain jumped for joy, and the fans cheered because we won the state championship." print(split_complex_sentence(test_sent2)) # 输出: ['The team captain jumped for joy', 'the fans cheered', 'because we won the state championship']
二、NLTK:基于句法树的分句提取
NLTK可借助Penn Treebank句法树,通过识别S(独立分句节点)和SBAR(从属分句节点)来拆分多分句句子。
实现逻辑
用NLTK的句法解析器生成树结构,递归遍历树节点,提取所有S和SBAR对应的文本内容。
代码示例
import nltk from nltk.tree import Tree # 下载必要依赖模型 nltk.download('punkt') nltk.download('averaged_perceptron_tagger') nltk.download('treebank') nltk.download('universal_tagset') def split_with_nltk(text): # 加载通用句法规则 parser = nltk.ChartParser(nltk.data.load('file:english_universal.cfg')) tokens = nltk.word_tokenize(text) clauses = [] def extract_clauses(node): if isinstance(node, Tree): # 提取分句节点对应的文本 if node.label() in ['S', 'SBAR']: clause = " ".join(node.leaves()).strip() clauses.append(clause.lstrip(',').strip()) # 递归遍历子节点 for child in node: extract_clauses(child) for tree in parser.parse(tokens): extract_clauses(tree) # 去重返回 return list(set(clauses)) # 测试 test_sent = "Though Mitchell prefers watching romantic films, he rented the latest spy thriller, and he enjoyed it very much." print(split_with_nltk(test_sent))
三、其他可选工具
如果需要更高精度的分句效果,可尝试:
- AllenNLP:提供专业的依存句法和 constituency 句法分析模块,对复杂句式的识别准确率更高。
- Hugging Face Transformers:使用预训练的句法分析模型(如
roberta-large-constituency-parser),可直接输出分句结构。
注意事项
- 无明显连词的隐含分句(仅靠标点分隔)需结合语义分析,但多数复杂句依赖句法标记(连词、从属引导词),因此句法分析仍是核心方案。
- 优先使用spaCy的
en_core_web_trf模型,其基于Transformer架构,对长难句的句法分析精度远高于基础模型。
内容的提问来源于stack exchange,提问作者DevPy
相关产品推荐
相关产品推荐

