You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在spaCy 3.5的NER预测中获取置信度分数?

获取spaCy 3.5自定义NER模型的预测置信度分数

spaCy默认不会在doc.ents中直接返回实体预测的置信度分数,但可以通过以下两种方式获取:

方法一:调用NER管道的predict方法手动解析分数

这种方式可以直接获取模型对每个token的标签预测概率,进而得到实体的置信度:

import spacy
from spacy.training import offsets_to_biluo_tags

# 加载自定义NER模型
nlp = spacy.load("your_custom_ner_model")
ner = nlp.get_pipe("ner")

# 测试文本
test_text = "需要预测的测试文本内容"

# 1. 创建未处理的Doc对象(不运行任何管道)
doc = nlp.make_doc(test_text)

# 2. 获取NER模型的原始预测分数
# scores是形状为 (token数量, 标签数量) 的数组,每个元素对应token对应标签的概率
scores = ner.predict(doc)

# 3. 获取已预测的实体,并转换为BILUO标签格式
predicted_ents = nlp(test_text).ents
biluo_tags = offsets_to_biluo_tags(doc, [(ent.start_char, ent.end_char, ent.label_) for ent in predicted_ents])

# 4. 解析每个实体的置信度
current_ent = []
current_label = None
current_scores = []
for token, tag, score in zip(doc, biluo_tags, scores):
    if tag.startswith("B-"):
        # 结束上一个实体(如果有)
        if current_ent:
            avg_score = sum(current_scores) / len(current_scores)
            print(f"实体: {' '.join([t.text for t in current_ent])}, 标签: {current_label}, 平均置信度: {avg_score:.4f}")
            current_ent = []
            current_scores = []
        # 开始新实体
        current_label = tag.split("-")[1]
        current_ent.append(token)
        label_id = ner.vocab.strings[current_label]
        current_scores = [score[label_id]]
    elif tag.startswith("I-"):
        # 继续当前实体
        current_ent.append(token)
        label_id = ner.vocab.strings[current_label]
        current_scores.append(score[label_id])
    elif tag == "O" and current_ent:
        # 结束当前实体
        avg_score = sum(current_scores) / len(current_scores)
        print(f"实体: {' '.join([t.text for t in current_ent])}, 标签: {current_label}, 平均置信度: {avg_score:.4f}")
        current_ent = []
        current_scores = []

# 处理最后一个实体
if current_ent:
    avg_score = sum(current_scores) / len(current_scores)
    print(f"实体: {' '.join([t.text for t in current_ent])}, 标签: {current_label}, 平均置信度: {avg_score:.4f}")

方法二:启用NER管道的输出分数直接获取

通过配置NER管道输出分数,可以直接从Token对象中获取对应标签的概率:

import spacy

# 加载自定义NER模型
nlp = spacy.load("your_custom_ner_model")
ner = nlp.get_pipe("ner")

# 启用NER管道的输出分数功能
ner.set_output(enable_output_scores=True)

test_text = "需要预测的测试文本内容"
doc = nlp(test_text)

for ent in doc.ents:
    # 实体的置信度可以取起始token对应标签的概率,或所有token的平均概率
    label_id = ner.vocab.strings[ent.label_]
    # 获取起始token的标签概率
    start_score = doc[ent.start]._.scores[label_id]
    # 计算实体所有token的平均概率
    avg_score = sum([doc[i]._.scores[label_id] for i in range(ent.start, ent.end)]) / (ent.end - ent.start)
    print(f"实体: {ent.text}, 标签: {ent.label_}, 起始token置信度: {start_score:.4f}, 平均置信度: {avg_score:.4f}")

说明

  • 实体的置信度可以根据需求选择起始token的标签概率(代表模型识别实体开头的置信度),或者实体所有对应token的平均概率(代表整个实体的整体置信度)。
  • 如果你的自定义模型是用spaCy标准流程训练的,以上两种方法都可以正常工作。

内容的提问来源于stack exchange,提问作者Ritaprava Dutta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 16:44:59