You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在spaCy中提取动词短语 解决ROOT子树提取失效问题

spaCy 主句动词短语提取方案

原有基于依存子树截取的方法仅适用于主语、宾语、状语从句这类本身对应独立完整依存子树的成分。主句核心动词的依存标签为ROOT,它的子树天然覆盖整句所有内容,直接截取会把主语、从句、宾语等无关成分全部包含,必须通过依存关系做过滤,才能精准提取动词短语。

实现逻辑

  • 遍历句法分析结果,定位依存标签为ROOT的主句核心谓词
  • 以核心谓词为中心,仅收集动词短语内部构成成分:助动词(含被动语态助动词)、修饰动词的副词、否定词,直接排除主语、宾语、状语从句、介词短语等非VP附属成分
  • 将收集到的所有成分按原文出现顺序排序后拼接,得到最终动词短语

代码实现

新增动词短语提取函数get_vp,完整可运行代码如下:

import spacy

def get_subj(decomp):
    for token in decomp:
        if ("subj" in token.dep_):
            subtree = list(token.subtree)
            start = subtree[0].i
            end = subtree[-1].i + 1
            return str(decomp[start:end])

def get_obj(decomp):
    for token in decomp:
        if ("dobj" in token.dep_ or "pobj" in token.dep_):
            subtree = list(token.subtree)
            start = subtree[0].i
            end = subtree[-1].i + 1
            return str(decomp[start:end])

def get_advcl(decomp):
    for token in decomp:
        if ("advcl" in token.dep_):
            subtree = list(token.subtree)
            start = subtree[0].i
            end = subtree[-1].i + 1
            return str(decomp[start:end])

def get_vp(decomp):
    # 定位主句ROOT核心谓词
    root = None
    for token in decomp:
        if token.dep_ == "ROOT":
            root = token
            break
    if not root:
        return ""
    vp_tokens = {root}
    # 定义动词短语内部允许的依存关系
    allowed_deps = {"aux", "aux:pass", "advmod", "neg"}
    for child in root.children:
        if child.dep_ in allowed_deps:
            vp_tokens.add(child)
            # 递归收集修饰助词/副词的附属成分,适配very/extremely等程度副词修饰场景
            for sub_child in child.children:
                if sub_child.dep_ in allowed_deps:
                    vp_tokens.add(sub_child)
    # 按原文顺序排序后截取
    vp_tokens = sorted(vp_tokens, key=lambda x: x.i)
    start = vp_tokens[0].i
    end = vp_tokens[-1].i + 1
    return str(decomp[start:end])

phrase = "Ultimate Swirly Ice Cream Scoopers are usually overrated when one considers all of the scoopers one could buy."

nlp = spacy.load("en_core_web_sm")
decomp = nlp(phrase)

subj = get_subj(decomp)
obj = get_obj(decomp)
advcl = get_advcl(decomp)
vp = get_vp(decomp)

print("subj: ", subj)
print("obj: ", obj)
print("advcl: ", advcl)
print("vp: ", vp)

运行结果

subj:  Ultimate Swirly Ice Cream Scoopers
obj:  all of the scoopers
advcl:  when one considers all of the scoopers one could buy
vp:  are usually overrated

扩展适配

如果需要覆盖更复杂的动词短语场景,可在allowed_deps集合中补充对应依存标签:

  • 短语动词小品词(如pick up中的up)添加prt标签
  • 并列动词结构添加cc、conj标签
  • 情态动词(can/should等)本身属于aux标签,无需额外配置

内容的提问来源于stack exchange,提问作者Chris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 20:39:19