直接调用HuggingFace模型与Token Classification Pipeline输出不一致问题
解决自定义调用NER模型与Pipeline输出不一致的问题
以下是导致输出差异的核心原因及修正方案:
核心差异点
- Tokenizer预处理参数不匹配
Pipeline默认使用padding="longest"(按批次最长文本填充),而非固定max_length;且会自动启用return_offsets_mapping=True来追踪原始文本与subword的映射关系,你的自定义代码缺少这些参数。 - 特殊Token未过滤
Pipeline会自动忽略[CLS]、[SEP]、[PAD]这类特殊token的标签结果,你的代码保留了这些无效位置的标签。 - Subword实体未合并
NER模型输出的是每个subword的标签,Pipeline会自动将属于同一实体的subword合并为单个实体条目,你的代码仅输出原始subword标签,未做合并处理。 - 标签映射未对齐
需要确保argmax得到的标签ID与模型的id2label映射对应,避免标签错位。
修正后的自定义代码
import torch from transformers import AutoTokenizer, AutoModelForTokenClassification # 加载模型和tokenizer ner_tokenizer = AutoTokenizer.from_pretrained("51la5/roberta-large-NER") model = AutoModelForTokenClassification.from_pretrained("51la5/roberta-large-NER") TORCH_DEVICE = "cuda" if torch.cuda.is_available() else "cpu" model.to(TORCH_DEVICE) text_batch = dataset['text'] # 匹配Pipeline的tokenizer参数 encodings_batch = ner_tokenizer( text_batch, padding="longest", # 替换为Pipeline默认的填充方式 truncation=True, return_tensors="pt", return_offsets_mapping=True # 必须添加,用于关联subword与原始文本 ) # 移到GPU input_ids = encodings_batch['input_ids'].to(TORCH_DEVICE) attention_mask = encodings_batch['attention_mask'].to(TORCH_DEVICE) # 模型推理(必须传入attention_mask,Pipeline会自动处理) with torch.no_grad(): outputs = model(input_ids, attention_mask=attention_mask)[0] # 获取标签ID并映射为标签文本 label_ids = outputs.argmax(dim=2).cpu().numpy() id2label = model.config.id2label # 处理每个文本的结果,对齐Pipeline输出格式 final_results = [] for i in range(len(text_batch)): offsets = encodings_batch['offset_mapping'][i].numpy() labels = label_ids[i] text = text_batch[i] entities = [] current_entity = None current_start = None current_label = None for idx, (offset, label_id) in enumerate(zip(offsets, labels)): # 跳过特殊token([CLS]、[SEP]、[PAD]) if offset[0] == 0 and offset[1] == 0: continue label = id2label[label_id] # 基于BIO标签体系处理实体合并 if label.startswith("B-"): # 保存上一个未结束的实体 if current_entity is not None: entities.append({ "entity": current_entity, "score": None, # 如需得分可对outputs做softmax后提取对应概率 "start": current_start, "end": offset[0], "entity_group": current_label }) # 启动新实体 current_entity = text[offset[0]:offset[1]] current_start = offset[0] current_label = label[2:] elif label.startswith("I-") and current_entity is not None: # 合并subword到当前实体 current_entity += text[offset[0]:offset[1]] else: # 非实体或实体结束,保存当前实体 if current_entity is not None: entities.append({ "entity": current_entity, "score": None, "start": current_start, "end": offset[0], "entity_group": current_label }) current_entity = None current_start = None current_label = None # 处理最后一个未结束的实体 if current_entity is not None: entities.append({ "entity": current_entity, "score": None, "start": current_start, "end": len(text), "entity_group": current_label }) final_results.append(entities) # final_results的格式与Pipeline输出完全一致
关键细节补充
- Attention Mask:原始代码未传入
attention_mask,模型会将padding token视为有效输入,导致标签错误,Pipeline会自动传入该参数。 - 得分计算:如果需要和Pipeline一样包含实体置信度得分,可对
outputs做softmax后提取对应标签的概率值。 - 标签体系适配:示例基于BIO标签体系,若模型使用BIOES等其他体系,需调整实体合并逻辑。
内容的提问来源于stack exchange,提问作者Bunnyrabbit
相关产品推荐
相关产品推荐

