You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何输出HuggingFace预训练Token分类模型的分类标签?

Token分类结果标签映射问题解决

你当前拿到的是模型输出的logits(未归一化的预测得分),还需要以下步骤才能得到和token对应的标签:

步骤1:从logits提取预测标签索引

logits的形状为[batch_size, sequence_length, num_labels],直接对每个token的logits取最大值的索引即可得到类别编号:

import torch as pt
from transformers import AutoTokenizer, AutoModelForTokenClassification

tokenizer = AutoTokenizer.from_pretrained("lwachowiak/Metaphor-Detection-XLMR")
model = AutoModelForTokenClassification.from_pretrained("lwachowiak/Metaphor-Detection-XLMR")

inputs = tokenizer("The virus is an enemy", return_tensors="pt")
outputs = model(**inputs)

# 获取每个token的预测标签索引
predictions = pt.argmax(outputs.logits, dim=-1)

步骤2:关联分词token与标签

通过tokenizer把输入id转回token,同时过滤掉模型的特殊token(如XLM-R的<s>、</s>):

# 转换为token列表
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
# 把张量形式的预测结果转为列表
predicted_label_ids = predictions[0].tolist()

步骤3:映射标签索引到实际标签

模型配置中自带id2label映射表,直接用它将索引转为对应标签:

# 遍历输出每个token和对应标签
for token, label_id in zip(tokens, predicted_label_ids):
    if token not in tokenizer.special_tokens_map.values():
        print(f"Token: {token}, 标签: {model.config.id2label[label_id]}")

执行后你会看到enemy对应的标签为1,符合预期结果。

内容的提问来源于stack exchange,提问作者Idkwhatywantmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 11:20:56