You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DistilBert二分类任务中标签形状不兼容问题及指标求解

修复DistilBert二分类任务中预测维度异常,正确计算召回率、精确率与AUC

问题根源

你的预测结果被意外展平成了(426530, 2),本质是test_dataset的维度处理错误——原本1930个样本,每个样本的token序列长度是221,数据集被错误地展平成了一维,导致模型预测时把每个token当成了独立样本,最终输出维度变成1930×221=426530个"伪样本"的预测结果。


分步解决方案

1. 修正测试数据集的批次结构

确保test_dataset按样本批次组织,而非展平的token序列。如果之前的数据集没有设置批次,添加batch()处理:

# 按实际显存调整batch_size,比如32、64等
test_dataset = test_dataset.batch(batch_size=32)

验证数据集形状:处理后,每个批次的input_ids形状应为(batch_size, 221),而非一维数组。

2. 生成维度匹配的预测结果

用修正后的数据集重新预测,确保输出形状为(1930, 2)(对应1930个样本的二分类概率):

import numpy as np

# 获取模型预测概率
predictions = model.predict(test_dataset, verbose=1)
# 转换为类别标签(取概率最大的类别)
rounded_predictions = np.argmax(predictions, axis=1)

# 若真实标签y_test_encoded是one-hot格式,转换为类别索引
y_test_labels = np.argmax(y_test_encoded, axis=1)
# 若y_test_encoded已经是类别索引(形状(1930,)),直接使用即可

3. 计算召回率、精确率与AUC

提供两种常用实现方式:

方式一:使用TensorFlow Keras指标
import tensorflow as tf

# 召回率
recall_metric = tf.keras.metrics.Recall()
recall_metric.update_state(y_test_labels, rounded_predictions)
recall = recall_metric.result().numpy()

# 精确率
precision_metric = tf.keras.metrics.Precision()
precision_metric.update_state(y_test_labels, rounded_predictions)
precision = precision_metric.result().numpy()

# AUC(需用正类的原始预测概率,不能用类别标签)
auc_metric = tf.keras.metrics.AUC(num_thresholds=100)
auc_metric.update_state(y_test_labels, predictions[:, 1])  # 取第二列作为正类概率
auc = auc_metric.result().numpy()

print(f"召回率: {recall:.4f}, 精确率: {precision:.4f}, AUC: {auc:.4f}")
方式二:使用Scikit-learn(更简洁)
from sklearn.metrics import recall_score, precision_score, roc_auc_score

recall = recall_score(y_test_labels, rounded_predictions)
precision = precision_score(y_test_labels, rounded_predictions)
# AUC使用正类的原始预测概率
auc = roc_auc_score(y_test_labels, predictions[:, 1])

print(f"召回率: {recall:.4f}, 精确率: {precision:.4f}, AUC: {auc:.4f}")

关键注意事项

  • 计算AUC时必须使用模型输出的原始概率值,不能用四舍五入后的类别标签,否则会严重影响指标准确性。
  • 确认真实标签和预测标签的维度完全匹配:均为(1930,)的类别索引,或均为(1930,2)的one-hot编码(此时计算指标时需对应处理)。

内容的提问来源于stack exchange,提问作者Injarapu Sri Sharanya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 15:05:23