基于BERT微调多标签分类模型时遇RuntimeError类型转换错误
多标签分类训练报错:Float无法转为Long的解决方法
问题场景
使用beomi/kcbert-base预训练模型微调多标签分类模型时,执行trainer.train()出现以下错误:
RuntimeError: result type Float can't be cast to the desired output type Long
错误原因
- 预处理函数错误:将批量样本的句子列表转为单个字符串传入tokenizer,导致输入格式异常,触发后续类型转换错误。
- 标签类别配置错误:标签列表包含了输入文本列
문장,导致模型输出维度与实际标签维度不匹配。 - 标签类型处理不当:直接返回
torch.tensor在数据集映射时可能被自动转换为numpy数组,类型处理不够稳妥。
修复步骤
1. 修正预处理函数
批量处理每个样本的句子,避免将整个批次转为单个字符串,并正确设置标签类型:
def preprocess_function(examples): # 批量处理句子,启用截断和填充适配模型输入长度 tokenized_examples = tokenizer(examples["문장"], truncation=True, padding=True) # 将标签转为float32类型的numpy数组,适配数据集映射逻辑 tokenized_examples['labels'] = np.array(examples["label"], dtype=np.float32) return tokenized_examples
2. 修正标签类别配置
移除标签列表中的输入文本列문장,确保num_labels与实际标签数量一致:
# 仅保留真实标签类别 labeling = ['식중독', '변질', '이취/냄새', '이물질', '배송', '포장', '가격불만', '품질', '유통기한', 'clean'] num_labels = len(labeling)
3. 确认模型配置正确性
确保模型初始化时num_labels和problem_type设置正确:
model = BertForSequenceClassification.from_pretrained( model_name, num_labels=num_labels, problem_type="multi_label_classification" ) model.config.id2label = {i: label for i, label in zip(range(num_labels), labeling)} model.config.label2id = {label: i for i, label in zip(range(num_labels), labeling)}
4. 验证compute_metrics函数
保留原函数即可,label_ranking_average_precision_score可兼容0/1整数类型的真实标签和浮点型预测值:
def compute_metrics(x): return { 'lrap': label_ranking_average_precision_score(x.label_ids, x.predictions), }
完成以上修改后,重新运行训练代码即可解决类型转换错误。
内容的提问来源于stack exchange,提问作者munirot
相关产品推荐
相关产品推荐

