如何用spaCy与textacy实现英语句子结构识别(SVO/SVOO等)
嘿,我来帮你搞定这个用spaCy和textacy识别英语句子结构的需求!下面是完整的实现方案,包含可运行的代码和详细说明:
使用spaCy和textacy识别英语句子结构
需求回顾
程序需读取段落,返回每个句子的SVO、SVOO、SVVO或自定义结构类型,示例参考:
- The cat sat on the mat - SVO
- The cat jumped and picked up the biscuit - SVVO
- The cat ate the biscuit and cookies - SVOO
完整实现代码
# -*- coding: utf-8 -*- #!/usr/bin/env python from __future__ import unicode_literals # 加载所需库 import en_core_web_sm import textacy # 加载spaCy英文轻量模型 nlp = en_core_web_sm.load() def identify_sentence_structure(sentence): doc = nlp(sentence) structure = "" # 提取主语(包括主动和被动主语) subjects = [token for token in doc if token.dep_ in ("nsubj", "nsubjpass")] # 提取核心谓语(ROOT节点动词,以及并列谓语) verbs = [token for token in doc if token.dep_ == "ROOT"] conj_verbs = [token for token in doc if token.dep_ == "conj" and token.head.pos_ == "VERB"] total_verbs = len(verbs) + len(conj_verbs) # 提取直接/间接宾语、表语(核心补语) objects = [token for token in doc if token.dep_ in ("dobj", "dative", "attr")] conj_objects = [token for token in doc if token.dep_ == "conj" and token.head.dep_ in ("dobj", "dative", "attr")] total_objects = len(objects) + len(conj_objects) # 匹配用户定义的结构规则 if total_verbs == 1 and total_objects == 1: structure = "SVO" elif total_verbs == 1 and total_objects >= 2: structure = "SVOO" elif total_verbs >= 2 and total_objects >= 1: structure = "SVVO" else: # 自定义其他结构类型,比如SVP(主系表)、SV(主谓)等 structure = "Custom (non-SVO/SVOO/SVVO)" return structure def analyze_paragraph(paragraph): doc = nlp(paragraph) results = [] # 遍历段落中的每个句子 for sent in doc.sents: sent_text = sent.text.strip() struct_type = identify_sentence_structure(sent_text) results.append(f"{sent_text} - {struct_type}") return results # 测试示例 if __name__ == "__main__": test_paragraph = """The cat sat on the mat. The cat jumped and picked up the biscuit. The cat ate the biscuit and cookies. She gave him a book and a pen. He runs fast.""" analysis_results = analyze_paragraph(test_paragraph) for res in analysis_results: print(res)
代码关键说明
- 模型选择:用
en_core_web_sm轻量模型,兼顾速度和基本句法分析能力,适合日常文本处理。 - 依存关系分析:通过spaCy的依存标签精准提取主语、谓语、宾语,同时处理并列结构(
conj标签),避免漏判并列的谓语或宾语。 - 结构匹配逻辑:完全贴合你给出的示例规则,同时预留了自定义结构的扩展空间,方便适配更多句式。
测试输出示例
运行代码后,测试段落会输出:
The cat sat on the mat - SVO The cat jumped and picked up the biscuit - SVVO The cat ate the biscuit and cookies - SVOO She gave him a book and a pen - SVOO He runs fast. - Custom (non-SVO/SVOO/SVVO)
内容的提问来源于stack exchange,提问作者Programmer_nltk
相关产品推荐
相关产品推荐

