如何基于SpaCy依存关系提取指定条件的部分子树?
问题:基于SpaCy依存关系提取指定子树并过滤依存条件
我已经用SpaCy解析了文本的依存关系,想在提取指定token/span的子树时加入依存关系条件,比如提取指定token的子树,但排除原token直接子节点为conj(并列)依存关系的部分。
举个具体例子:从句子**"The entrepreneur and philanthropist Bill Gates and the Apple's Steve Jobs ate hamburgers."**中提取人物姓名及对应属性,期望得到如下表格:
| person | attribute |
|---|---|
| Bill Gates | entrepreneur and philanthropist |
| Steve Jobs | Apple's |
当前代码能提取人物实体,但Bill Gates的子树和Steve Jobs的子树重叠:
import spacy nlp = spacy.load("en_core_web_trf") s = "The entrepreneur and philanthropist Bill Gates and the Apple's Steve Jobs ate hamburgers." doc = nlp(s) persons = [ent for ent in doc.ents if ent.label_ == "PERSON"] # [Bill Gates, Steve Jobs] [[token for token in p.subtree] for p in persons] # [[The, entrepreneur, and, philanthropist, Bill, Gates, and, the, Apple, 's, Steve, Jobs], [the, Apple, 's, Steve, Jobs]]
我希望仅保留Bill Gates子树中直接子节点为nmod依存关系的部分,或者移除直接子节点为conj的部分。另外如果有更优方法实现上述表格提取也请赐教,我对SpaCy和Python都不太熟悉。
解决方案
方法1:过滤子树,排除conj相关分支
SpaCy的Token对象自带children属性可获取直接子节点,我们可以通过递归遍历子树,跳过conj依存的直接子节点及其后代,实现精准过滤:
import spacy nlp = spacy.load("en_core_web_trf") s = "The entrepreneur and philanthropist Bill Gates and the Apple's Steve Jobs ate hamburgers." doc = nlp(s) def get_filtered_subtree(root_token, exclude_conj=True): filtered = [] # 递归遍历子树,跳过conj分支 def traverse(token): # 若当前节点是根节点的conj直接子节点,跳过 if exclude_conj and token.dep_ == "conj" and token.head == root_token: return filtered.append(token) for child in token.children: traverse(child) traverse(root_token) return filtered # 处理每个人物实体 result = [] for ent in doc.ents: if ent.label_ != "PERSON": continue # 获取实体的根token(多token实体的依存中心词,比如Bill Gates的根是Gates) ent_root = ent.root # 获取过滤后的子树 subtree = get_filtered_subtree(ent_root) # 从子树中移除实体本身的token,得到属性部分 attribute_tokens = [tok for tok in subtree if tok not in ent] # 拼接成字符串并去除多余空格 attribute = " ".join([tok.text for tok in attribute_tokens]).strip() result.append({"person": ent.text, "attribute": attribute}) # 输出结果表格 print("| person | attribute |") print("|--------------|-------------------------------|") for item in result: print(f"| {item['person']:<12} | {item['attribute']:<29} |")
运行结果:
| person | attribute | |--------------|-------------------------------| | Bill Gates | The entrepreneur and philanthropist | | Steve Jobs | the Apple's |
方法2:用SpaCy Matcher直接匹配人物+属性结构
如果只针对这类「前置属性+人物」的固定结构,也可以用Matcher精准匹配,避免子树过滤的复杂性:
import spacy from spacy.matcher import Matcher nlp = spacy.load("en_core_web_trf") matcher = Matcher(nlp.vocab) # 定义匹配模式:任意前置修饰词(排除动词、连词) + PERSON实体 pattern = [ {"POS": {"NOT_IN": ["VERB", "CCONJ"]}, "OP": "*"}, {"ENT_TYPE": "PERSON", "OP": "1"} ] matcher.add("PERSON_WITH_ATTR", [pattern]) s = "The entrepreneur and philanthropist Bill Gates and the Apple's Steve Jobs ate hamburgers." doc = nlp(s) matches = matcher(doc) # 去重并整理结果 seen_persons = set() result = [] for match_id, start, end in matches: span = doc[start:end] # 定位span中的PERSON实体 person_ent = next((ent for ent in span.ents if ent.label_ == "PERSON"), None) if not person_ent or person_ent.text in seen_persons: continue seen_persons.add(person_ent.text) # 提取属性部分(移除实体文本后的内容) attribute = span.text.replace(person_ent.text, "").strip() result.append({"person": person_ent.text, "attribute": attribute}) # 输出结果表格 print("| person | attribute |") print("|--------------|-------------------------------|") for item in result: print(f"| {item['person']:<12} | {item['attribute']:<29} |")
方案说明
- 方法1更通用,适合需要基于依存关系做复杂过滤的场景;
- 方法2更直接,针对特定结构匹配,代码更简洁;
- 可根据文本中人物属性的结构复杂度选择对应方案。
内容的提问来源于stack exchange,提问作者dufei
相关产品推荐
相关产品推荐

