如何使用Python库spaCy检测句子的主动与被动语态?
用spaCy检测英文句子的主动/被动语态
要判断英文句子的语态,核心是利用spaCy的依存句法分析功能:
- 主动语态的核心谓语动词,其主语的依存关系标记为
nsubj(普通主语) - 被动语态的核心谓语动词,其主语的依存关系标记为
nsubjpass(被动主语)
具体实现步骤
- 先安装spaCy和英文模型
pip install spacy python -m spacy download en_core_web_sm
- 编写检测函数,遍历句子中的动词判断语态
import spacy nlp = spacy.load("en_core_web_sm") def detect_voice(sentence): doc = nlp(sentence) for token in doc: # 检查被动主语标记或被动助动词(如was/were) if token.dep_ == "nsubjpass" or token.dep_ == "auxpass": return "被动语态" return "主动语态" # 测试示例句子 passive_sentence = "John was accused of committing crimes by David" active_sentence = "David accused John of committing crimes" print(detect_voice(passive_sentence)) # 输出:被动语态 print(detect_voice(active_sentence)) # 输出:主动语态
提取核心语态结构(对应示例需求)
如果需要像示例那样提取核心的语态结构,可以扩展函数:
def extract_core_voice_structure(sentence): doc = nlp(sentence) for token in doc: if token.dep_ == "nsubjpass": # 提取被动结构:主语 + 被动助动词 + 核心动词 core_parts = [token.text] root_verb = token.head # 查找被动助动词(如was) for child in root_verb.children: if child.dep_ == "auxpass": core_parts.append(child.text) core_parts.append(root_verb.text) return " ".join(core_parts) elif token.dep_ == "nsubj": # 提取主动结构:主语 + 核心动词 return f"{token.text} {token.head.text}" return "" print(extract_core_voice_structure(passive_sentence)) # 输出:John was accused print(extract_core_voice_structure(active_sentence)) # 输出:David accused John
内容的提问来源于stack exchange,提问作者sruthi
相关产品推荐
相关产品推荐

