You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于PyTorch与Hugging Face BERT的二分类改造出现KeyError问题

问题解决:KeyError与二分类模型适配

一、KeyError: 2047 错误原因及修复

错误根源

DataLoader迭代时会按连续整数索引(0,1,2,...)请求数据,但你的df_train/df_val索引可能不连续(比如筛选、删除操作导致索引断裂),当DataLoader请求索引2047时,DataFrame中没有对应标签的行,因此抛出KeyError。

修复方案

修改Dataset类,改用**位置索引(iloc)**获取标签,或提前重置数据框索引:

方案1:修改Dataset类

class Dataset(torch.utils.data.Dataset):
    
    def __init__(self, df):
        # 重置索引确保从0开始连续
        self.df = df.reset_index(drop=True)
        self.labels = self.df['target']
        self.texts = [tokenizer(text, padding='max_length', max_length=512, truncation=True, return_tensors="pt") 
                      for text in self.df['text']]

    def classes(self):
        return self.labels

    def __len__(self):
        return len(self.labels)

    def get_batch_labels(self, idx):
        # 用iloc按位置取,避免标签索引不连续问题
        return np.array(self.labels.iloc[idx])

    def get_batch_texts(self, idx):
        return self.texts[idx]

    def __getitem__(self, idx):
        batch_texts = self.get_batch_texts(idx)
        batch_y = self.get_batch_labels(idx)
        return batch_texts, batch_y

方案2:创建Dataset前重置索引

df_train = df_train.reset_index(drop=True)
df_val = df_val.reset_index(drop=True)

二、二分类模型的损失函数与准确率计算修正

1. 损失函数错误

nn.CrossEntropyLoss()是为多分类任务设计的,二分类需改用:

  • 模型带sigmoid时用nn.BCELoss()
  • 模型不带sigmoid时用nn.BCEWithLogitsLoss()(更推荐,数值稳定性更好)

推荐修改(BCEWithLogitsLoss)

修改模型移除最后一层sigmoid:

class BertClassifier(nn.Module):
    def __init__(self, dropout=0.5):
        super(BertClassifier, self).__init__()
        self.bert = BertModel.from_pretrained('bert-base-cased')
        self.dropout = nn.Dropout(dropout)
        self.linear = nn.Linear(768, 1)

    def forward(self, input_id, mask):
        _, pooled_output = self.bert(input_ids=input_id, attention_mask=mask, return_dict=False)
        dropout_output = self.dropout(pooled_output)
        linear_output = self.linear(dropout_output)
        return linear_output  # 直接返回logits,无需sigmoid

训练函数中替换损失函数:

criterion = nn.BCEWithLogitsLoss()

2. 准确率计算错误

output.argmax(dim=1)是多分类逻辑,二分类单输出需用阈值判断:

# 训练阶段修正
output = model(input_id, mask)
train_label = train_label.unsqueeze(1).float()
batch_loss = criterion(output, train_label)
# 用sigmoid转概率,0.5作为分类阈值
preds = torch.sigmoid(output) >= 0.5
acc = (preds == train_label).sum().item()
total_acc_train += acc

# 验证阶段同理
output = model(input_id, mask)
val_label = val_label.unsqueeze(1).float()
batch_loss = criterion(output, val_label)
preds = torch.sigmoid(output) >= 0.5
acc = (preds == val_label).sum().item()
total_acc_val += acc

三、其他细节优化

  1. 设备迁移:取消注释CUDA迁移代码,确保模型、数据在同一设备:
if use_cuda:
    model = model.cuda()
    criterion = criterion.cuda()

同时将输入和标签移到对应设备:

mask = train_input['attention_mask'].to(device)
input_id = train_input['input_ids'].squeeze(1).to(device)
train_label = train_label.unsqueeze(1).float().to(device)
  1. Tokenizer输出简化:提前处理tokenizer输出,避免重复squeeze:
# 修改Dataset初始化
self.texts = [tokenizer(text, padding='max_length', max_length=512, truncation=True, return_tensors="pt")['input_ids'].squeeze(0) 
              for text in self.df['text']]
self.attn_masks = [tokenizer(text, padding='max_length', max_length=512, truncation=True, return_tensors="pt")['attention_mask'].squeeze(0) 
                   for text in self.df['text']]

内容的提问来源于stack exchange,提问作者Aniket Gaudgaul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 14:01:42