You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于SpaCy依存关系提取指定条件的部分子树?

问题:基于SpaCy依存关系提取指定子树并过滤依存条件

我已经用SpaCy解析了文本的依存关系,想在提取指定token/span的子树时加入依存关系条件,比如提取指定token的子树,但排除原token直接子节点为conj(并列)依存关系的部分。

举个具体例子:从句子**"The entrepreneur and philanthropist Bill Gates and the Apple's Steve Jobs ate hamburgers."**中提取人物姓名及对应属性,期望得到如下表格:

personattribute
Bill Gatesentrepreneur and philanthropist
Steve JobsApple's

当前代码能提取人物实体,但Bill Gates的子树和Steve Jobs的子树重叠:

import spacy
nlp = spacy.load("en_core_web_trf")

s = "The entrepreneur and philanthropist Bill Gates and the Apple's Steve Jobs ate hamburgers."
doc = nlp(s)

persons = [ent for ent in doc.ents if ent.label_ == "PERSON"]
# [Bill Gates, Steve Jobs]

[[token for token in p.subtree] for p in persons]
# [[The, entrepreneur, and, philanthropist, Bill, Gates, and, the, Apple, 's, Steve, Jobs], [the, Apple, 's, Steve, Jobs]]

我希望仅保留Bill Gates子树中直接子节点为nmod依存关系的部分,或者移除直接子节点为conj的部分。另外如果有更优方法实现上述表格提取也请赐教,我对SpaCy和Python都不太熟悉。


解决方案

方法1:过滤子树,排除conj相关分支

SpaCy的Token对象自带children属性可获取直接子节点,我们可以通过递归遍历子树,跳过conj依存的直接子节点及其后代,实现精准过滤:

import spacy
nlp = spacy.load("en_core_web_trf")

s = "The entrepreneur and philanthropist Bill Gates and the Apple's Steve Jobs ate hamburgers."
doc = nlp(s)

def get_filtered_subtree(root_token, exclude_conj=True):
    filtered = []
    # 递归遍历子树,跳过conj分支
    def traverse(token):
        # 若当前节点是根节点的conj直接子节点,跳过
        if exclude_conj and token.dep_ == "conj" and token.head == root_token:
            return
        filtered.append(token)
        for child in token.children:
            traverse(child)
    traverse(root_token)
    return filtered

# 处理每个人物实体
result = []
for ent in doc.ents:
    if ent.label_ != "PERSON":
        continue
    # 获取实体的根token(多token实体的依存中心词,比如Bill Gates的根是Gates)
    ent_root = ent.root
    # 获取过滤后的子树
    subtree = get_filtered_subtree(ent_root)
    # 从子树中移除实体本身的token,得到属性部分
    attribute_tokens = [tok for tok in subtree if tok not in ent]
    # 拼接成字符串并去除多余空格
    attribute = " ".join([tok.text for tok in attribute_tokens]).strip()
    result.append({"person": ent.text, "attribute": attribute})

# 输出结果表格
print("| person       | attribute                     |")
print("|--------------|-------------------------------|")
for item in result:
    print(f"| {item['person']:<12} | {item['attribute']:<29} |")

运行结果:

| person       | attribute                     |
|--------------|-------------------------------|
| Bill Gates   | The entrepreneur and philanthropist |
| Steve Jobs   | the Apple's                   |

方法2:用SpaCy Matcher直接匹配人物+属性结构

如果只针对这类「前置属性+人物」的固定结构,也可以用Matcher精准匹配,避免子树过滤的复杂性:

import spacy
from spacy.matcher import Matcher

nlp = spacy.load("en_core_web_trf")
matcher = Matcher(nlp.vocab)

# 定义匹配模式:任意前置修饰词(排除动词、连词) + PERSON实体
pattern = [
    {"POS": {"NOT_IN": ["VERB", "CCONJ"]}, "OP": "*"},
    {"ENT_TYPE": "PERSON", "OP": "1"}
]
matcher.add("PERSON_WITH_ATTR", [pattern])

s = "The entrepreneur and philanthropist Bill Gates and the Apple's Steve Jobs ate hamburgers."
doc = nlp(s)

matches = matcher(doc)
# 去重并整理结果
seen_persons = set()
result = []
for match_id, start, end in matches:
    span = doc[start:end]
    # 定位span中的PERSON实体
    person_ent = next((ent for ent in span.ents if ent.label_ == "PERSON"), None)
    if not person_ent or person_ent.text in seen_persons:
        continue
    seen_persons.add(person_ent.text)
    # 提取属性部分(移除实体文本后的内容)
    attribute = span.text.replace(person_ent.text, "").strip()
    result.append({"person": person_ent.text, "attribute": attribute})

# 输出结果表格
print("| person       | attribute                     |")
print("|--------------|-------------------------------|")
for item in result:
    print(f"| {item['person']:<12} | {item['attribute']:<29} |")

方案说明

  • 方法1更通用,适合需要基于依存关系做复杂过滤的场景;
  • 方法2更直接,针对特定结构匹配,代码更简洁;
  • 可根据文本中人物属性的结构复杂度选择对应方案。

内容的提问来源于stack exchange,提问作者dufei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 23:07:36