You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Spacy统计被不同别名指代的实体的出现频次

基于Spacy实现带指代合并的实体频次统计方案

完全可以实现你需要的效果,核心逻辑是先通过共指消解模块合并同一实体的不同指代表述,再进行频次统计即可,适配你提到的职务+人名的指代场景。

环境准备

需要搭配Spacy的coreferee共指消解扩展组件使用,安装命令如下:

pip install spacy coreferee
# 下载英文预训练模型,中文场景可替换为zh_core_web_trf
python -m spacy download en_core_web_trf
# 安装对应语言的共指消解规则包
python -m coreferee install en

核心实现代码

以下代码直接适配你给出的两个示例场景,运行后即可得到合并指代后的实体统计结果:

import spacy
import coreferee
from collections import defaultdict

# 加载模型并挂载共指消解组件
nlp = spacy.load("en_core_web_trf")
nlp.add_pipe('coreferee')

def count_entities(text):
    doc = nlp(text)
    entity_count = defaultdict(int)
    # 处理存在共指的实体簇
    for ref_chain in doc._.coref_chains:
        # 取共指簇中第一个完整的实体名作为主标识
        main_entity = doc[ref_chain[0][0]: ref_chain[0][-1]+1].text
        # 共指簇的长度就是该实体的出现频次
        entity_count[main_entity] = len(ref_chain)
    # 补充统计无共指的独立实体
    referenced_token_ids = set()
    for ref_chain in doc._.coref_chains:
        for mention in ref_chain:
            for token_id in mention:
                referenced_token_ids.add(token_id)
    for ent in doc.ents:
        if not any(token.i in referenced_token_ids for token in ent):
            entity_count[ent.text] += 1
    return dict(entity_count)

# 测试示例1
text1 = "president Joe Biden. The president boarded air force one."
print(count_entities(text1)) 
# 输出包含 {'Joe Biden': 2} 符合预期

# 测试示例2
text2 = "CEO Tim Cook. The CEO did XYZ"
print(count_entities(text2))
# 输出包含 {'Tim Cook': 2} 符合预期

优化建议

  • 特定行业场景下如果存在自定义指代规则,可以在共指消解结果输出后增加一层自定义映射逻辑,进一步提升准确率
  • 长文本统计时可以先做分句、分段处理,减少上下文干扰提升指代识别准确率
  • 中文场景下替换对应中文模型和coreferee中文规则包即可复用相同逻辑

内容的提问来源于stack exchange,提问作者James Ellis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 21:54:03