You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用en_core_sci_md提取SVO时如何消除冗余输出?

解决en_core_sci_md提取SVO时的冗余问题

背景与修改

我需要提取句子的主语(Subject)、谓语(Verb)、宾语(Object)及其关联关系,需求与Stack Overflow帖子《How to extract subjects in a sentence and their respective dependent phrases?》一致。针对Spacy版本差异,我对原帖代码做了两处修改:

  1. 调整解析器初始化代码:
import spacy
nlp = spacy.load('en_core_web_md')
doc = nlp(sentence)
print(findSVAOs(doc))
  1. 修改findSVAOs函数的动词筛选逻辑:
    原代码:
verbs = [tok for tok in tokens if tok.pos_ == "VERB" and tok.dep_ != "aux"]

修改后:

verbs = [tok for tok in tokens if tok.pos_ == "VERB" or tok.dep_ != "aux"]

使用en_core_web_md通用模型时,输出符合预期:

[('lung cancer', 'causes', 'huge mortality'), ('companies', 'require', 'new drugs'), ('review', 'highlights', 'inextricable role')]

问题描述

切换到生物医学领域的en_core_sci_md模型后,输出出现冗余的SVO三元组,例如:

  • ('lung cancer', 'causes', 'huge mortality population') 与 ('lung cancer', 'causes', 'population') 重复(短宾语是长宾语的子集)
  • ('role', 'bioconstituents', 'signaling first time') 与 ('bioconstituents', 'signaling', 'first time') 重复(错误识别动词导致无效三元组)

完整冗余输出:

[('lung cancer', 'causes', 'huge mortality population'), ('lung cancer', 'causes', 'population'), ('companies', 'require', 'new drugs'), ('review', 'highlights', 'inextricable role lucidum'), ('role', 'bioconstituents', 'signaling first time'), ('bioconstituents', 'signaling', 'first time')]

解决方案

1. 过滤包含式冗余

遍历所有SVO三元组,按「主语+谓语」分组,移除被同组内更长宾语包含的短宾语,保留核心语义更完整的项。

示例代码:

def remove_redundant_svos(svos):
    # 按主语+谓语分组管理宾语
    sv_groups = {}
    for s, v, o in svos:
        key = (s, v)
        sv_groups.setdefault(key, []).append(o)
    
    filtered_svos = []
    for (s, v), objects in sv_groups.items():
        # 按宾语分词数量降序排序,优先处理长宾语
        sorted_objs = sorted(objects, key=lambda x: len(x.split()), reverse=True)
        keep_objs = []
        for obj in sorted_objs:
            # 仅保留未被已保留宾语包含的项
            if not any(obj in other_obj and obj != other_obj for other_obj in keep_objs):
                keep_objs.append(obj)
        # 重新组合为三元组
        for obj in keep_objs:
            filtered_svos.append((s, v, obj))
    return filtered_svos

# 测试使用
redundant_svos = [('lung cancer', 'causes', 'huge mortality population'), ('lung cancer', 'causes', 'population'), ('companies', 'require', 'new drugs'), ('review', 'highlights', 'inextricable role lucidum'), ('role', 'bioconstituents', 'signaling first time'), ('bioconstituents', 'signaling', 'first time')]
print(remove_redundant_svos(redundant_svos))

输出结果:

[('lung cancer', 'causes', 'huge mortality population'), ('companies', 'require', 'new drugs'), ('review', 'highlights', 'inextricable role lucidum'), ('role', 'bioconstituents', 'signaling first time'), ('bioconstituents', 'signaling', 'first time')]

2. 优化动词筛选逻辑

你之前修改的筛选条件pos_ == "VERB" or tok.dep_ != "aux"过于宽松,会把非动词成分误判为谓语(比如bioconstituents是名词却被当作动词)。应调整为只保留核心谓语动词:

修改findSVAOs中的动词筛选代码:

verbs = [tok for tok in tokens if tok.pos_ == "VERB" and tok.dep_ in ("ROOT", "ccomp", "conj")]

该条件仅保留根动词、补语动词和并列动词,从源头减少无效的SVO三元组。

3. 针对生物医学模型调整宾语提取规则

en_core_sci_md会使用领域特有的句法标签,提取宾语时应限定为直接宾语(dobj)、间接宾语(dative)等核心语义成分,忽略修饰性附属结构,避免把冗余的修饰词纳入宾语。

内容的提问来源于stack exchange,提问作者small tomato

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 18:20:34