You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

直接调用HuggingFace模型与Token Classification Pipeline输出不一致问题

解决自定义调用NER模型与Pipeline输出不一致的问题

以下是导致输出差异的核心原因及修正方案:

核心差异点

  1. Tokenizer预处理参数不匹配
    Pipeline默认使用padding="longest"(按批次最长文本填充),而非固定max_length;且会自动启用return_offsets_mapping=True来追踪原始文本与subword的映射关系,你的自定义代码缺少这些参数。
  2. 特殊Token未过滤
    Pipeline会自动忽略[CLS]、[SEP]、[PAD]这类特殊token的标签结果,你的代码保留了这些无效位置的标签。
  3. Subword实体未合并
    NER模型输出的是每个subword的标签,Pipeline会自动将属于同一实体的subword合并为单个实体条目,你的代码仅输出原始subword标签,未做合并处理。
  4. 标签映射未对齐
    需要确保argmax得到的标签ID与模型的id2label映射对应,避免标签错位。

修正后的自定义代码

import torch
from transformers import AutoTokenizer, AutoModelForTokenClassification

# 加载模型和tokenizer
ner_tokenizer = AutoTokenizer.from_pretrained("51la5/roberta-large-NER")
model = AutoModelForTokenClassification.from_pretrained("51la5/roberta-large-NER")
TORCH_DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
model.to(TORCH_DEVICE)

text_batch = dataset['text']

# 匹配Pipeline的tokenizer参数
encodings_batch = ner_tokenizer(
    text_batch,
    padding="longest",  # 替换为Pipeline默认的填充方式
    truncation=True,
    return_tensors="pt",
    return_offsets_mapping=True  # 必须添加,用于关联subword与原始文本
)

# 移到GPU
input_ids = encodings_batch['input_ids'].to(TORCH_DEVICE)
attention_mask = encodings_batch['attention_mask'].to(TORCH_DEVICE)

# 模型推理(必须传入attention_mask,Pipeline会自动处理)
with torch.no_grad():
    outputs = model(input_ids, attention_mask=attention_mask)[0]

# 获取标签ID并映射为标签文本
label_ids = outputs.argmax(dim=2).cpu().numpy()
id2label = model.config.id2label

# 处理每个文本的结果,对齐Pipeline输出格式
final_results = []
for i in range(len(text_batch)):
    offsets = encodings_batch['offset_mapping'][i].numpy()
    labels = label_ids[i]
    text = text_batch[i]
    entities = []
    
    current_entity = None
    current_start = None
    current_label = None
    
    for idx, (offset, label_id) in enumerate(zip(offsets, labels)):
        # 跳过特殊token([CLS]、[SEP]、[PAD])
        if offset[0] == 0 and offset[1] == 0:
            continue
        
        label = id2label[label_id]
        # 基于BIO标签体系处理实体合并
        if label.startswith("B-"):
            # 保存上一个未结束的实体
            if current_entity is not None:
                entities.append({
                    "entity": current_entity,
                    "score": None,  # 如需得分可对outputs做softmax后提取对应概率
                    "start": current_start,
                    "end": offset[0],
                    "entity_group": current_label
                })
            # 启动新实体
            current_entity = text[offset[0]:offset[1]]
            current_start = offset[0]
            current_label = label[2:]
        elif label.startswith("I-") and current_entity is not None:
            # 合并subword到当前实体
            current_entity += text[offset[0]:offset[1]]
        else:
            # 非实体或实体结束,保存当前实体
            if current_entity is not None:
                entities.append({
                    "entity": current_entity,
                    "score": None,
                    "start": current_start,
                    "end": offset[0],
                    "entity_group": current_label
                })
                current_entity = None
                current_start = None
                current_label = None
    # 处理最后一个未结束的实体
    if current_entity is not None:
        entities.append({
            "entity": current_entity,
            "score": None,
            "start": current_start,
            "end": len(text),
            "entity_group": current_label
        })
    final_results.append(entities)

# final_results的格式与Pipeline输出完全一致

关键细节补充

  • Attention Mask:原始代码未传入attention_mask,模型会将padding token视为有效输入,导致标签错误,Pipeline会自动传入该参数。
  • 得分计算:如果需要和Pipeline一样包含实体置信度得分,可对outputs做softmax后提取对应标签的概率值。
  • 标签体系适配:示例基于BIO标签体系,若模型使用BIOES等其他体系,需调整实体合并逻辑。

内容的提问来源于stack exchange,提问作者Bunnyrabbit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 12:40:42