如何通过spaCy的Scorer/Example类提取TP/FP/TN/FN并计算混淆矩阵
问题:从spaCy的NER评分结果中提取TP、FP、TN、FN指标
我尝试使用spaCy API计算命名实体识别(NER)模型的准确率(Accuracy)和特异度(Specificity)。scorer.score_spans(example)方法可计算模型预测实体的精确率(Precision)、召回率(Recall)和F1值,但无法直接获取真阳性(TP)、假阳性(FP)、真阴性(TN)、假阴性(FN)。
当前代码与数据结构
评分代码
import spacy from spacy.scorer import Scorer from spacy.training.example import Example scorer = Scorer() example = [] for obs in example_list: print('Input for a prediction:', obs['full_text']) pred = custom_nlp(obs['full_text']) ## custom_nlp是自定义的NER模型 print('Predicted based off of input:', pred, '// Entities being reviewed:', obs['entities']) temp = Example.from_dict(pred, {'entities': obs['entities']}) example.append(temp) scores = scorer.score_spans(example, "ents")
输入数据结构示例
example_list[0] {'full_text': 'I would like to remove my kid Florence from the will. How do I do that?', 'entities': [(30, 38, 'PERSON')]}
现有评分结果
{'ents_p': 0.8731019522776573, 'ents_r': 0.9179019384264538, 'ents_f': 0.8949416342412452, 'ents_per_type': {'PERSON': {'p': 0.9039145907473309, 'r': 0.9694656488549618, 'f': 0.9355432780847145}, 'GPE': {'p': 0.7973856209150327, 'r': 0.9384615384615385, 'f': 0.8621908127208481}, 'STREET_ADDRESS': {'p': 0.8308457711442786, 'r': 0.893048128342246, 'f': 0.8608247422680412}, 'ORGANIZATION': {'p': 0.9565217391304348, 'r': 0.7415730337078652, 'f': 0.8354430379746837}, 'CREDIT_CARD': {'p': 0.9411764705882353, 'r': 1.0, 'f': 0.9696969696969697}, 'AGE': {'p': 1.0, 'r': 1.0, 'f': 1.0}, 'US_SSN': {'p': 1.0, 'r': 1.0, 'f': 1.0}, 'DOMAIN_NAME': {'p': 0.4, 'r': 1.0, 'f': 0.5714285714285715}, 'TITLE': {'p': 0.8709677419354839, 'r': 0.84375, 'f': 0.8571428571428571}, 'PHONE_NUMBER': {'p': 0.8275862068965517, 'r': 0.8275862068965517, 'f': 0.8275862068965517}, 'EMAIL_ADDRESS': {'p': 1.0, 'r': 1.0, 'f': 1.0}, 'DATE_TIME': {'p': 1.0, 'r': 1.0, 'f': 1.0}, 'NRP': {'p': 1.0, 'r': 1.0, 'f': 1.0}, 'IBAN_CODE': {'p': 1.0, 'r': 1.0, 'f': 1.0}, 'IP_ADDRESS': {'p': 0.75, 'r': 0.75, 'f': 0.75}, 'ZIP_CODE': {'p': 0.8333333333333334, 'r': 0.7142857142857143, 'f': 0.7692307692307692}, 'US_DRIVER_LICENSE': {'p': 1.0, 'r': 1.0, 'f': 1.0}}}
解决方案
spaCy的scorer.score_spans没有直接返回TP/FP/FN/TN,但可以通过手动遍历Example统计(最可靠)或现有指标反推(局限性大)获取。
方法1:手动遍历Example统计TP/FP/FN
直接对比每个样本的标注实体与预测实体,基于严格匹配规则(实体起始/结束位置、类型完全一致)统计:
def count_tp_fp_fn(examples): tp = 0 fp = 0 fn = 0 for example in examples: # 将标注实体转换为集合,方便交集/差集计算 gold_ents = set((ent.start_char, ent.end_char, ent.label_) for ent in example.reference.ents) # 将预测实体转换为集合 pred_ents = set((ent.start_char, ent.end_char, ent.label_) for ent in example.predicted.ents) # 真阳性:同时存在于标注和预测中的实体 tp += len(gold_ents & pred_ents) # 假阳性:预测存在但标注不存在的实体 fp += len(pred_ents - gold_ents) # 假阴性:标注存在但预测不存在的实体 fn += len(gold_ents - pred_ents) return tp, fp, fn # 使用你的example列表调用函数 tp, fp, fn = count_tp_fp_fn(example) print(f"TP: {tp}, FP: {fp}, FN: {fn}")
方法2:计算TN(真阴性)
NER中的TN指非实体文本被正确预测为非实体,由于文本中非实体跨度数量极大,通常不直接统计。若需计算,可按token级别统计:
def count_tn(examples): tn = 0 for example in examples: doc = example.predicted # 收集标注实体覆盖的token索引 gold_token_indices = set() for ent in example.reference.ents: for idx in range(ent.start, ent.end): gold_token_indices.add(idx) # 收集预测实体覆盖的token索引 pred_token_indices = set() for ent in example.predicted.ents: for idx in range(ent.start, ent.end): pred_token_indices.add(idx) # 统计既不在标注也不在预测中的token数量,即为TN total_tokens = len(doc) tn += total_tokens - len(gold_token_indices | pred_token_indices) return tn tn = count_tn(example) print(f"TN: {tn}")
计算准确率与特异度
得到TP/FP/FN/TN后,即可计算所需指标:
- 准确率:
accuracy = (tp + tn) / (tp + fp + fn + tn) - 特异度:
specificity = tn / (tn + fp)
方法3:通过现有P/R/F1反推(局限性大)
若仅需整体的TP/FP/FN,可利用精确率(P)、召回率(R)的公式反推,但需要先知道总标注实体数或总预测实体数:
- 精确率:
P = TP / (TP + FP) - 召回率:
R = TP / (TP + FN) - F1值:
F1 = 2*P*R/(P+R)
但由于spaCy未直接返回总标注/预测实体数,这种方法不如手动统计可靠,仅适用于快速估算。
内容的提问来源于stack exchange,提问作者lucaslaff3
相关产品推荐
相关产品推荐

