SpaCy模型返回空字典无预测结果的技术求助
问题:SpaCy SpanCat模型返回空字典预测结果
模型创建与测试流程
- 使用Prodigy以spans方式标注数据:
python -m prodigy spans.manual test_model_1 blank:en test_data.txt --label label1,label2,label3
- 训练模型:
python -m prodigy train ./test_model_1 --spancat test_model_1
- 测试代码返回空字典:
import spacy from spacy.tokens import Doc # Define a function to get all possible predictions for a document def get_all_cats(doc): return {cat: doc._.cats[cat].cats.items() for cat in doc._.cats} # Load the model nlp = spacy.load("test_model_1/model-best") # Add the 'cats' and 'all_cats' extensions to the Doc class Doc.set_extension("cats", default={}, force=True) Doc.set_extension("all_cats", getter=get_all_cats) # The text to analyze text = "This is a sample sentence to get predictions on." # Process the text with the model doc = nlp(text) # Get all possible predictions for the document predictions = doc._.all_cats # Print the predictions print(predictions)
环境信息
- spaCy版本:3.5.1
- 平台:Windows-10
- Python版本:3.10.4
- en_core_web_sm版本:3.5.0
问题原因
混淆了SpanCat(跨度分类)和TextCat(文本分类)的模型输出属性。doc._.cats是文本分类模型的结果存储位置,而你训练的是跨度分类模型,预测结果不会存放在这个属性里,因此返回空字典。
解决方案
修改测试代码,获取SpanCat模型的正确预测结果:
方法1:直接获取分组跨度结果
doc.spans返回按标签分组的跨度结果,默认组名为"sc":
import spacy # 加载模型 nlp = spacy.load("test_model_1/model-best") # 测试文本 text = "This is a sample sentence to get predictions on." # 处理文本 doc = nlp(text) # 获取默认组的跨度预测(训练时自定义组名需替换对应名称) spans_predictions = doc.spans.get("sc", []) # 打印每个跨度的细节 for span in spans_predictions: print(f"文本: {span.text}, 标签: {span.label_}, 置信度: {span.score:.4f}")
方法2:查看所有候选跨度及评分
如果需要完整的候选跨度和对应置信度,可通过doc._.spancat获取:
import spacy nlp = spacy.load("test_model_1/model-best") text = "This is a sample sentence to get predictions on." doc = nlp(text) # 获取SpanCat组件 spancat = nlp.get_pipe("spancat") # 获取候选跨度和对应评分 candidates = doc._.spancat.candidates scores = doc._.spancat.scores # 遍历输出所有候选信息 for candidate, score in zip(candidates, scores): best_label = spancat.labels[score.argmax()] span_text = doc[candidate.start:candidate.end].text print(f"跨度: {span_text}, 最优标签: {best_label}, 置信度: {score.max():.4f}")
额外检查项
- 确认标注数据包含有效跨度,若标注为空或无匹配测试文本的内容,模型也会返回空结果。
- 若训练时通过
--spancat-group指定了自定义组名,需替换doc.spans中的"sc"为对应名称。 - 查看训练日志,确认模型训练过程中损失下降、指标提升,确保模型学到有效特征。
内容的提问来源于stack exchange,提问作者MrAutoIt
相关产品推荐
相关产品推荐

