如何修改1D CNN模型以避免依赖零填充序列的填充长度?
问题背景
我正在基于时序数据开展特征预测任务,使用如下1D CNN模型:
import torch import torch.nn as nn import torch.nn.functional as F class My1DCNN(nn.Module): def __init__(self): super(My1DCNN, self).__init__() self.conv1 = nn.Conv1d(in_channels=32, out_channels=64, kernel_size=5, stride=1, padding=1) self.conv2 = nn.Conv1d(in_channels=64, out_channels=64, kernel_size=5, stride=1, padding=1) self.conv3 = nn.Conv1d(in_channels=64, out_channels=128, kernel_size=5, stride=1, padding=1) self.conv4 = nn.Conv1d(in_channels=128, out_channels=128, kernel_size=5, stride=1, padding=1) self.conv5 = nn.Conv1d(in_channels=128, out_channels=256, kernel_size=5, stride=1, padding=1) self.maxpool = nn.MaxPool1d(kernel_size=2, stride=2) self.fc1 = nn.Linear(1280, 239 * 4) # Adjust the input features to the correct flattened size def forward(self, x, eeg_mask=None): if eeg_mask is not None: eeg_mask = eeg_mask.unsqueeze(1) x = x * eeg_mask x = F.relu(self.conv1(x)) x = self.maxpool(x) x = F.relu(self.conv2(x)) x = self.maxpool(x) x = F.relu(self.conv3(x)) x = self.maxpool(x) x = F.relu(self.conv4(x)) x = self.maxpool(x) x = F.relu(self.conv5(x)) x = self.maxpool(x) x = x.view(x.size(0), -1) # Flatten keeping the batch size intact x = self.fc1(x) x = x.view(-1, 239, 4) return x
该模型接收形状为(batch, channel, time)的批量序列,输出形状为(batch, time, feature)的预测特征。为处理变长序列,采用零填充并使用eeg_mask掩码标记填充部分。
模型在训练和测试集上表现良好,但输入torch.rand_like(x)生成的随机数据时,模型仍能取得不错的表现。这表明模型将零填充部分的长度作为特征使用,不符合预期——模型应仅依赖真实数据而非零值填充长度。
已尝试方案
- 结合
pack_padded_sequence()的GRU
在CNN前加入GRU,使用pack_padded_sequence()和pad_packed_sequence()处理掩码,但为适配CNN必须指定total_length=239,再次引入零填充,导致后续CNN层仍学习填充长度的模式,未解决问题。尝试的代码如下:class My1DCNN(nn.Module): def __init__(self, input_size=32, hidden_size=32, num_layers=1): super(My1DCNN, self).__init__() # Set up 1D convolutional layers self.gru = nn.GRU(input_size, hidden_size, num_layers, batch_first=True) self.conv1 = nn.Conv1d( in_channels=32, out_channels=64, kernel_size=5, stride=1, padding=1 ) self.conv2 = nn.Conv1d( in_channels=64, out_channels=64, kernel_size=5, stride=1, padding=1 ) self.conv3 = nn.Conv1d( in_channels=64, out_channels=128, kernel_size=5, stride=1, padding=1 ) self.conv4 = nn.Conv1d( in_channels=128, out_channels=128, kernel_size=5, stride=1, padding=1 ) self.conv5 = nn.Conv1d( in_channels=128, out_channels=256, kernel_size=5, stride=1, padding=1 ) # Max pooling for 1D self.maxpool = nn.MaxPool1d(kernel_size=2, stride=2) # Assuming the size after conv and pooling layers, adjust accordingly self.fc1 = nn.Linear( 1280, 239 * 4 ) # Adjust the input features to the correct flattened size def forward(self, x, eeg_mask=None): # Calculate the lengths of the sequences based on the eeg_mask if eeg_mask is not None: lengths = eeg_mask.sum(dim=1).cpu().int() x = x * eeg_mask.unsqueeze(1) else: lengths = torch.full((x.size(0),), x.size(1), dtype=torch.int64) x = x.permute(0, 2, 1) # Pack the padded sequence packed_input = pack_padded_sequence( x, lengths, batch_first=True, enforce_sorted=False ) packed_output, _ = self.gru(packed_input) gru_output, _ = pad_packed_sequence( packed_output, batch_first=True, total_length=239 ) gru_output = gru_output.permute(0, 2, 1) x = F.relu(self.conv1(gru_output)) x = self.maxpool(x) x = F.relu(self.conv2(x)) x = self.maxpool(x) x = F.relu(self.conv3(x)) x = self.maxpool(x) x = F.relu(self.conv4(x)) x = self.maxpool(x) x = F.relu(self.conv5(x)) x = self.maxpool(x) # Flatten the output for the fully connected layer x = x.view(x.size(0), -1) # Flatten keeping the batch size intact x = self.fc1(x) # Reshape if needed, or adjust fc layer output size x = x.view(-1, 239, 4) return x - 镜像填充替代零填充
替换填充方式后,模型仍会学习填充带来的虚假模式,无效果。
可行修改方案
1. 卷积层后动态掩码,阻断填充信息传递
在每一层卷积激活后、池化前,根据原始eeg_mask生成对应缩放后的掩码,将填充区域的特征置为极小值(如-1e9),让MaxPool1d完全忽略这些区域的信息,避免池化层捕捉到填充相关模式。
示例修改的forward函数:
def forward(self, x, eeg_mask=None): batch_size = x.size(0) current_mask = eeg_mask # 初始掩码 shape (batch, time) # 第一层卷积+掩码+池化 x = F.relu(self.conv1(x)) if current_mask is not None: # 对应conv1(kernel=5,padding=1)的有效区域:时间维度裁剪为原长度-2 current_mask = current_mask[:, 2:-2].unsqueeze(1) # shape (batch,1,time_conv) x = torch.where(current_mask == 0, torch.tensor(-1e9, device=x.device), x) x = self.maxpool(x) if current_mask is not None: # 池化后时间维度减半,掩码同步采样 current_mask = current_mask[:, :, ::2] # 第二层卷积+掩码+池化 x = F.relu(self.conv2(x)) if current_mask is not None: current_mask = current_mask[:, :, 2:-2] x = torch.where(current_mask == 0, torch.tensor(-1e9, device=x.device), x) x = self.maxpool(x) if current_mask is not None: current_mask = current_mask[:, :, ::2] # 重复逻辑处理conv3-conv5 x = F.relu(self.conv3(x)) if current_mask is not None: current_mask = current_mask[:, :, 2:-2] x = torch.where(current_mask == 0, torch.tensor(-1e9, device=x.device), x) x = self.maxpool(x) if current_mask is not None: current_mask = current_mask[:, :, ::2] x = F.relu(self.conv4(x)) if current_mask is not None: current_mask = current_mask[:, :, 2:-2] x = torch.where(current_mask == 0, torch.tensor(-1e9, device=x.device), x) x = self.maxpool(x) if current_mask is not None: current_mask = current_mask[:, :, ::2] x = F.relu(self.conv5(x)) if current_mask is not None: current_mask = current_mask[:, :, 2:-2] x = torch.where(current_mask == 0, torch.tensor(-1e9, device=x.device), x) x = self.maxpool(x) # 后续全连接层处理 x = x.view(batch_size, -1) x = self.fc1(x) x = x.view(-1, 239, 4) return x
注意:需根据卷积的kernel、padding、stride准确计算掩码的裁剪范围,确保掩码与特征图时间维度完全匹配。
2. 自适应池化替代固定最大池化
将MaxPool1d替换为AdaptiveMaxPool1d,指定输出时间维度大小,无论输入有效序列长度如何,都能得到一致的输出维度。同时结合掩码在池化前置填充区域特征为极小值,确保自适应池化仅基于有效区域计算。
3. 移除全连接层固定维度依赖,改用全局池化+动态映射
原模型fc1的输入维度固定为1280,依赖固定输入长度和池化次数。可将最后一层卷积后的特征用全局平均池化(或全局最大池化)得到样本固定维度特征,再通过全连接层映射到输出特征维度,让模型仅关注有效区域特征,不再依赖填充长度。
修改示例:
class My1DCNN(nn.Module): def __init__(self): super(My1DCNN, self).__init__() self.conv1 = nn.Conv1d(in_channels=32, out_channels=64, kernel_size=5, stride=1, padding=1) self.conv2 = nn.Conv1d(in_channels=64, out_channels=64, kernel_size=5, stride=1, padding=1) self.conv3 = nn.Conv1d(in_channels=64, out_channels=128, kernel_size=5, stride=1, padding=1) self.conv4 = nn.Conv1d(in_channels=128, out_channels=128, kernel_size=5, stride=1, padding=1) self.conv5 = nn.Conv1d(in_channels=128, out_channels=256, kernel_size=5, stride=1, padding=1) self.maxpool = nn.MaxPool1d(kernel_size=2, stride=2) self.global_avg_pool = nn.AdaptiveAvgPool1d(1) # 全局平均池化到1个时间步 self.fc1 = nn.Linear(256, 4) # 映射到目标特征维度 def forward(self, x, eeg_mask=None): if eeg_mask is not None: eeg_mask = eeg_mask.unsqueeze(1) x = x * eeg_mask x = F.relu(self.conv1(x)) x = self.maxpool(x) x = F.relu(self.conv2(x)) x = self.maxpool(x) x = F.relu(self.conv3(x)) x = self.maxpool(x) x = F.relu(self.conv4(x)) x = self.maxpool(x) x = F.relu(self.conv5(x)) # 全局池化提取核心特征 x = self.global_avg_pool(x).squeeze(-1) # shape (batch, 256) x = self.fc1(x) # shape (batch,4) # 扩展到目标时间维度239 x = x.unsqueeze(1).repeat(1,239,1) # shape (batch,239,4) return x
4. 训练时加入对抗性正则化
训练过程中,随机打乱不同样本的填充长度,或生成不同填充长度的随机输入,计算模型在这些输入上的损失并加入总损失,惩罚模型依赖填充长度的行为。比如每次迭代随机选取部分样本,替换为随机数据但保持不同填充长度,迫使模型放弃对填充长度的依赖。
内容的提问来源于stack exchange,提问作者Kim Junyoung

