使用en_core_sci_md提取SVO时如何消除冗余输出?
解决en_core_sci_md提取SVO时的冗余问题
背景与修改
我需要提取句子的主语(Subject)、谓语(Verb)、宾语(Object)及其关联关系,需求与Stack Overflow帖子《How to extract subjects in a sentence and their respective dependent phrases?》一致。针对Spacy版本差异,我对原帖代码做了两处修改:
- 调整解析器初始化代码:
import spacy nlp = spacy.load('en_core_web_md') doc = nlp(sentence) print(findSVAOs(doc))
- 修改
findSVAOs函数的动词筛选逻辑:
原代码:
verbs = [tok for tok in tokens if tok.pos_ == "VERB" and tok.dep_ != "aux"]
修改后:
verbs = [tok for tok in tokens if tok.pos_ == "VERB" or tok.dep_ != "aux"]
使用en_core_web_md通用模型时,输出符合预期:
[('lung cancer', 'causes', 'huge mortality'), ('companies', 'require', 'new drugs'), ('review', 'highlights', 'inextricable role')]
问题描述
切换到生物医学领域的en_core_sci_md模型后,输出出现冗余的SVO三元组,例如:
('lung cancer', 'causes', 'huge mortality population')与('lung cancer', 'causes', 'population')重复(短宾语是长宾语的子集)('role', 'bioconstituents', 'signaling first time')与('bioconstituents', 'signaling', 'first time')重复(错误识别动词导致无效三元组)
完整冗余输出:
[('lung cancer', 'causes', 'huge mortality population'), ('lung cancer', 'causes', 'population'), ('companies', 'require', 'new drugs'), ('review', 'highlights', 'inextricable role lucidum'), ('role', 'bioconstituents', 'signaling first time'), ('bioconstituents', 'signaling', 'first time')]
解决方案
1. 过滤包含式冗余
遍历所有SVO三元组,按「主语+谓语」分组,移除被同组内更长宾语包含的短宾语,保留核心语义更完整的项。
示例代码:
def remove_redundant_svos(svos): # 按主语+谓语分组管理宾语 sv_groups = {} for s, v, o in svos: key = (s, v) sv_groups.setdefault(key, []).append(o) filtered_svos = [] for (s, v), objects in sv_groups.items(): # 按宾语分词数量降序排序,优先处理长宾语 sorted_objs = sorted(objects, key=lambda x: len(x.split()), reverse=True) keep_objs = [] for obj in sorted_objs: # 仅保留未被已保留宾语包含的项 if not any(obj in other_obj and obj != other_obj for other_obj in keep_objs): keep_objs.append(obj) # 重新组合为三元组 for obj in keep_objs: filtered_svos.append((s, v, obj)) return filtered_svos # 测试使用 redundant_svos = [('lung cancer', 'causes', 'huge mortality population'), ('lung cancer', 'causes', 'population'), ('companies', 'require', 'new drugs'), ('review', 'highlights', 'inextricable role lucidum'), ('role', 'bioconstituents', 'signaling first time'), ('bioconstituents', 'signaling', 'first time')] print(remove_redundant_svos(redundant_svos))
输出结果:
[('lung cancer', 'causes', 'huge mortality population'), ('companies', 'require', 'new drugs'), ('review', 'highlights', 'inextricable role lucidum'), ('role', 'bioconstituents', 'signaling first time'), ('bioconstituents', 'signaling', 'first time')]
2. 优化动词筛选逻辑
你之前修改的筛选条件pos_ == "VERB" or tok.dep_ != "aux"过于宽松,会把非动词成分误判为谓语(比如bioconstituents是名词却被当作动词)。应调整为只保留核心谓语动词:
修改findSVAOs中的动词筛选代码:
verbs = [tok for tok in tokens if tok.pos_ == "VERB" and tok.dep_ in ("ROOT", "ccomp", "conj")]
该条件仅保留根动词、补语动词和并列动词,从源头减少无效的SVO三元组。
3. 针对生物医学模型调整宾语提取规则
en_core_sci_md会使用领域特有的句法标签,提取宾语时应限定为直接宾语(dobj)、间接宾语(dative)等核心语义成分,忽略修饰性附属结构,避免把冗余的修饰词纳入宾语。
内容的提问来源于stack exchange,提问作者small tomato
相关产品推荐
相关产品推荐

