在BERT模型上叠加CNN层微调时遇conv2d参数错误求助
问题分析与解决方案
核心错误原因
你遇到的conv2d()参数类型不匹配错误,直接诱因是nn.Conv2d中使用了无效的padding=True参数——PyTorch的Conv2d不支持布尔值作为padding参数(仅接受整数、元组或'valid'/'same'字符串),这会导致内部参数处理异常,引发类型不匹配的报错。
此外,代码中还存在两处潜在问题,会导致后续运行错误:
forward函数未返回损失值,但training_step试图解构loss, outputs- 全连接层
self.fc的输入维度计算错误,与卷积、池化后的张量形状不匹配
修复步骤
1. 修正Conv2d的padding参数
将padding=True替换为合法取值:
- 若使用PyTorch 1.10+,可直接用
padding='same'实现“输出尺寸与输入一致”的填充 - 或手动计算填充值:针对
kernel_size=(3, 768),使用padding=(1, 383)(对应seq_len维度填充1,hidden_size维度填充383)
2. 修正forward函数的返回逻辑
在forward中添加损失计算逻辑,当传入labels时返回(loss, predictions),否则仅返回预测结果,匹配training_step的解构需求。
3. 修正全连接层的输入维度
根据卷积、池化后的张量形状重新计算输入维度:
- 经过
Conv2d(kernel_size=(3,768), padding='same')后,张量形状为[batch_size, 13, 128, 1](hidden_size维度被卷积核覆盖,输出宽度为1) - 经过
MaxPool2d(kernel_size=3, stride=1)后,seq_len维度变为128 - 3 + 1 = 126,张量形状为[batch_size,13,126,1] - 展平后尺寸为
13 * 126 = 1638,因此全连接层应改为nn.Linear(1638, 3)
4. 确保设备一致性(可选)
确认模型与输入张量运行在同一设备(CPU/GPU),避免因设备不匹配引发的隐性错误。
修正后的完整代码
from transformers import BertModel, BertTokenizer, AdamW import torch.nn as nn import pytorch_lightning as pl import torch import torch.nn.functional as F class CommentModule(pl.LightningModule): def __init__(self, n_classes=3): super().__init__() self.bert = BertModel.from_pretrained('bert-base-uncased') # 修正padding参数为合法取值 self.conv = nn.Conv2d(in_channels=13, out_channels=13, kernel_size=(3, 768), padding='same') self.relu = nn.ReLU() self.pool = nn.MaxPool2d(kernel_size=3, stride=1) self.dropout = nn.Dropout(0.1) # 修正全连接层输入维度 self.fc = nn.Linear(13 * 126, 3) self.flat = nn.Flatten() self.softmax = nn.LogSoftmax(dim=1) def forward(self, input_ids, attention_mask, labels=None): # 获取BERT的所有隐藏层输出 outputs = self.bert(input_ids, attention_mask, output_hidden_states=True) # 调整维度:[batch_size, num_hidden_layers, seq_len, hidden_size] x = torch.transpose(torch.cat([t.unsqueeze(0) for t in outputs.hidden_states], 0), 0, 1) x = self.dropout(x) x = self.conv(x) x = self.relu(x) x = self.dropout(x) x = self.pool(x) # 展平后传入全连接层 x = self.flat(x) x = self.dropout(x) x = self.fc(x) predictions = self.softmax(x) # 计算损失(当labels存在时) if labels is not None: loss = F.nll_loss(predictions, labels) return loss, predictions return predictions def training_step(self, batch, batch_idx): input_ids = batch['input_ids'] attention_mask = batch['attention_mask'] labels = batch['labels'] loss, outputs = self.forward(input_ids, attention_mask, labels) self.log('train_loss', loss) return {'loss': loss, 'predictions': outputs, 'labels': labels}
内容的提问来源于stack exchange,提问作者trell
相关产品推荐
相关产品推荐

