Spacy遇E966错误无法训练NER模型,该如何正确训练?
解决Spacy NER训练中的[E966]错误
原始训练代码
nlp = spacy.blank('en') def train_model(train_data): if 'ner' not in nlp.pipe_names: ner = nlp.create_pipe('ner') nlp.add_pipe(ner, last = True) for _, annotation in train_data: for ent in annotation['entities']: ner.add_label(ent[2]) other_pipes = [pipe for pipe in nlp.pipe_names if pipe != 'ner'] with nlp.disable_pipes(*other_pipes): # only train NER optimizer = nlp.begin_training() for itn in range(10): print("Statring iteration " + str(itn)) random.shuffle(train_data) losses = {} index = 0 for text, annotations in train_data: try: nlp.update( [text], [annotations], drop=0.2, sgd=optimizer, losses=losses) except Exception as e: pass print(losses)
调用代码:
train_model(train_data)
错误信息
[E966]
nlp.add_pipe现在需要传入已注册组件工厂的字符串名称,而非可调用组件。预期为字符串,但得到<spacy.pipeline.ner.EntityRecognizer object at 0x7ff8f26c3a50>(名称:'None')。
- 如果您使用
nlp.create_pipe('name')创建组件:请移除nlp.create_pipe,改为调用nlp.add_pipe('name')。- 如果您传入了类似
TextCategorizer()的组件:请改用字符串名称调用nlp.add_pipe,例如nlp.add_pipe('textcat')。- 如果您使用自定义组件:请为自定义组件添加装饰器
@Language.component(适用于函数组件)或@Language.factory(适用于类组件/工厂)并指定名称,例如@Language.component('your_name')。之后您可以运行nlp.add_pipe('your_name')将其添加到流水线中。
修改后的代码
import random import spacy nlp = spacy.blank('en') def train_model(train_data): if 'ner' not in nlp.pipe_names: # 直接通过字符串名称添加NER组件,无需create_pipe nlp.add_pipe('ner', last=True) # 获取NER组件实例来添加实体标签 ner = nlp.get_pipe('ner') for _, annotation in train_data: for ent in annotation['entities']: ner.add_label(ent[2]) other_pipes = [pipe for pipe in nlp.pipe_names if pipe != 'ner'] with nlp.disable_pipes(*other_pipes): # 仅训练NER组件 optimizer = nlp.begin_training() for itn in range(10): print(f"开始迭代 {itn}") random.shuffle(train_data) losses = {} for text, annotations in train_data: try: nlp.update( [text], [annotations], drop=0.2, sgd=optimizer, losses=losses) except Exception as e: # 可选:打印错误信息便于调试 # print(f"处理文本失败: {text}, 错误详情: {e}") pass print(losses) train_model(train_data)
关键修改说明
- 移除
nlp.create_pipe:Spacy 3.x及以上版本要求nlp.add_pipe直接传入组件的字符串名称(如'ner'),不再支持传入组件实例。 - 获取NER组件实例:通过
nlp.get_pipe('ner')获取已添加的NER组件,用于后续添加实体标签。 - 优化细节:将字符串拼接改为f-string提升可读性;可选开启错误打印,方便定位训练中出现问题的样本。
内容的提问来源于stack exchange,提问作者slaam
相关产品推荐
相关产品推荐

