自定义NER模型无实体返回、迭代为空且报E109错误,求解决
自定义NER模型训练异常排查与解决
问题现象
- 训练时迭代输出为空字典,无loss更新:
Iteration Number:0 {} Iteration Number:1 {} Iteration Number:2 {}
- 测试时抛出初始化错误:
ValueError: [E109] Component 'ner_pipe' could not be run. Did you forget to call
initialize()?
- 此前正常训练时的实体识别输出:
Entities [('McVeggie', 'FoodProduct')] Entities [('McEgg', 'FoodProduct')] Entities [('McChicken', 'FoodProduct')] Entities [('McSpicy Paneer', 'FoodProduct')] Entities [('McSpicy Chicken', 'FoodProduct')]
训练代码
Training_data = [ ('what is the price of McVeggie?', {'entities': [(21, 29, 'FoodProduct')]}), ('what is the price of McEgg?', {'entities': [(21, 26, 'FoodProduct')]}), ('what is the price of McChicken?', {'entities': [(21, 30, 'FoodProduct')]}), ('what is the price of McSpicy Paneer?', {'entities': [(21, 35, 'FoodProduct')]}), ('what is the price of McSpicy Chicken?', {'entities': [(21, 36, 'FoodProduct')]}), ] # Testing sample data testing_sample='what is the price of McAloo?' import random from pathlib import Path import spacy from tqdm import tqdm from spacy.training.example import Example model = None output_dir=Path("C:\Users\Desktop\Folder1") iterations = 20 # loading a blank model if model is not None: nlp = spacy.load(model) print("Loaded model '%s'" % model) else: nlp = spacy.blank('en') print("Created blank 'en' model") # Setting up the pipeline if 'ner' not in nlp.pipe_names: ner_pipe = nlp.add_pipe('ner',name='ner_pipe',last=True) else: ner_pipe = nlp.get_pipe('ner') # Adding entities labels to the ner pipeline for text, annotations in Training_data: for entity in annotations.get('entities'): ner_pipe.add_label(entity[2]) # Getting names of other pipes to disable them during training other_pipes = [pipe for pipe in nlp.pipe_names if pipe != 'ner'] other_pipes # Training NER model with nlp.disable_pipes(*other_pipes): optimizer = nlp.begin_training() for itn in range(iterations): print("Iteration Number:" + str(itn)) random.shuffle(Training_data) losses = {} for text, annotations in Training_data: doc = nlp.make_doc(text) example = Example.from_dict(doc, annotations) nlp.update([example], drop=0.2, sgd=optimizer, losses=losses) print(losses)
问题原因及解决方法
1. 核心问题:NER组件未初始化
spaCy v3.x版本中,使用空白模型添加NER组件后,必须调用nlp.initialize()完成组件参数初始化,否则:
- 训练时无法更新模型参数,导致loss始终为空字典
- 测试时组件未就绪,抛出
[E109]初始化错误
修复方法:在添加完实体标签后(即ner_pipe.add_label()循环后)添加初始化代码:
# 初始化nlp管道 nlp.initialize()
2. 训练流程补全
- 路径转义问题:Windows路径需避免转义字符,将
output_dir改为原始字符串:output_dir=Path(r"C:\Users\Desktop\Folder1") - 保存训练后的模型:训练完成后保存模型,避免每次运行都从头训练,在训练循环结束后添加:
if output_dir is not None: output_dir.mkdir(parents=True, exist_ok=True) nlp.to_disk(output_dir) print(f"模型已保存至 {output_dir}") - 添加测试逻辑:训练后直接用
nlp对象处理测试样本,或加载保存的模型进行测试:# 测试训练后的模型 doc = nlp(testing_sample) entities = [(ent.text, ent.label_) for ent in doc.ents] print(f"Entities {entities}")
3. 提升模型稳定性
- 当前训练样本仅5条,数据量过少,建议补充更多同类型样本,避免模型过拟合或不稳定
- 调整
drop参数:小数据集下drop=0.2的 dropout 可能导致训练不稳定,可暂时设为0.1或移除该参数
内容的提问来源于stack exchange,提问作者Popeye
相关产品推荐
相关产品推荐

