You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SpaCy模型返回空字典无预测结果的技术求助

问题:SpaCy SpanCat模型返回空字典预测结果

模型创建与测试流程

  1. 使用Prodigy以spans方式标注数据:
python -m prodigy spans.manual test_model_1 blank:en test_data.txt --label label1,label2,label3
  1. 训练模型:
python -m prodigy train ./test_model_1 --spancat test_model_1
  1. 测试代码返回空字典:
import spacy
from spacy.tokens import Doc

# Define a function to get all possible predictions for a document
def get_all_cats(doc):
    return {cat: doc._.cats[cat].cats.items() for cat in doc._.cats}

# Load the model
nlp = spacy.load("test_model_1/model-best")

# Add the 'cats' and 'all_cats' extensions to the Doc class
Doc.set_extension("cats", default={}, force=True)
Doc.set_extension("all_cats", getter=get_all_cats)

# The text to analyze
text = "This is a sample sentence to get predictions on."

# Process the text with the model
doc = nlp(text)

# Get all possible predictions for the document
predictions = doc._.all_cats

# Print the predictions
print(predictions)

环境信息

  • spaCy版本:3.5.1
  • 平台:Windows-10
  • Python版本:3.10.4
  • en_core_web_sm版本:3.5.0

问题原因

混淆了SpanCat(跨度分类)和TextCat(文本分类)的模型输出属性。doc._.cats是文本分类模型的结果存储位置,而你训练的是跨度分类模型,预测结果不会存放在这个属性里,因此返回空字典。

解决方案

修改测试代码,获取SpanCat模型的正确预测结果:

方法1:直接获取分组跨度结果

doc.spans返回按标签分组的跨度结果,默认组名为"sc":

import spacy

# 加载模型
nlp = spacy.load("test_model_1/model-best")

# 测试文本
text = "This is a sample sentence to get predictions on."

# 处理文本
doc = nlp(text)

# 获取默认组的跨度预测(训练时自定义组名需替换对应名称)
spans_predictions = doc.spans.get("sc", [])

# 打印每个跨度的细节
for span in spans_predictions:
    print(f"文本: {span.text}, 标签: {span.label_}, 置信度: {span.score:.4f}")

方法2:查看所有候选跨度及评分

如果需要完整的候选跨度和对应置信度,可通过doc._.spancat获取:

import spacy

nlp = spacy.load("test_model_1/model-best")
text = "This is a sample sentence to get predictions on."
doc = nlp(text)

# 获取SpanCat组件
spancat = nlp.get_pipe("spancat")
# 获取候选跨度和对应评分
candidates = doc._.spancat.candidates
scores = doc._.spancat.scores

# 遍历输出所有候选信息
for candidate, score in zip(candidates, scores):
    best_label = spancat.labels[score.argmax()]
    span_text = doc[candidate.start:candidate.end].text
    print(f"跨度: {span_text}, 最优标签: {best_label}, 置信度: {score.max():.4f}")

额外检查项

  1. 确认标注数据包含有效跨度,若标注为空或无匹配测试文本的内容,模型也会返回空结果。
  2. 若训练时通过--spancat-group指定了自定义组名,需替换doc.spans中的"sc"为对应名称。
  3. 查看训练日志,确认模型训练过程中损失下降、指标提升,确保模型学到有效特征。

内容的提问来源于stack exchange,提问作者MrAutoIt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 22:35:22