You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于BERT微调多标签分类模型时遇RuntimeError类型转换错误

多标签分类训练报错:Float无法转为Long的解决方法

问题场景

使用beomi/kcbert-base预训练模型微调多标签分类模型时,执行trainer.train()出现以下错误:

RuntimeError: result type Float can't be cast to the desired output type Long

错误原因

  1. 预处理函数错误:将批量样本的句子列表转为单个字符串传入tokenizer,导致输入格式异常,触发后续类型转换错误。
  2. 标签类别配置错误:标签列表包含了输入文本列문장,导致模型输出维度与实际标签维度不匹配。
  3. 标签类型处理不当:直接返回torch.tensor在数据集映射时可能被自动转换为numpy数组,类型处理不够稳妥。

修复步骤

1. 修正预处理函数

批量处理每个样本的句子,避免将整个批次转为单个字符串,并正确设置标签类型:

def preprocess_function(examples):
    # 批量处理句子,启用截断和填充适配模型输入长度
    tokenized_examples = tokenizer(examples["문장"], truncation=True, padding=True)
    # 将标签转为float32类型的numpy数组,适配数据集映射逻辑
    tokenized_examples['labels'] = np.array(examples["label"], dtype=np.float32)
    return tokenized_examples

2. 修正标签类别配置

移除标签列表中的输入文本列문장,确保num_labels与实际标签数量一致:

# 仅保留真实标签类别
labeling = ['식중독', '변질', '이취/냄새', '이물질', '배송', '포장', '가격불만', '품질', '유통기한', 'clean']
num_labels = len(labeling)

3. 确认模型配置正确性

确保模型初始化时num_labels和problem_type设置正确:

model = BertForSequenceClassification.from_pretrained(
    model_name, 
    num_labels=num_labels, 
    problem_type="multi_label_classification"
)
model.config.id2label = {i: label for i, label in zip(range(num_labels), labeling)}
model.config.label2id = {label: i for i, label in zip(range(num_labels), labeling)}

4. 验证compute_metrics函数

保留原函数即可,label_ranking_average_precision_score可兼容0/1整数类型的真实标签和浮点型预测值:

def compute_metrics(x):
    return {
        'lrap': label_ranking_average_precision_score(x.label_ids, x.predictions),
    }

完成以上修改后,重新运行训练代码即可解决类型转换错误。

内容的提问来源于stack exchange,提问作者munirot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 04:16:04