Autoencoder训练遭遇隐形数值边界问题,求解决思路
Autoencoder拟合MinMax缩放股票数据的边界失真问题
问题描述
使用经过MinMax缩放的股票数据训练Autoencoder,模型由tanh激活函数与全连接层构建,当前latent_dim与输入层维度一致(后续计划缩减维度)。但模型在转换0.15以下和0.8以上的数值时存在明显失真,仿佛有隐形边界,无法实现输入输出1:1的拟合目标。
模型代码
编码器
class SparseEncoder(nn.Module): def __init__(self, input_shape: int, latent_dims, dtype=torch.float64): super().__init__() self.linear1 = nn.Linear(input_shape, 512, dtype=dtype) self.linear2 = nn.Linear(512, 256, dtype=dtype) self.linear3 = nn.Linear(256, 128, dtype=dtype) self.linear4 = nn.Linear(128, 64, dtype=dtype) self.linear5 = nn.Linear(64, 32, dtype=dtype) self.linear6 = nn.Linear(32, 16, dtype=dtype) self.linear7 = nn.Linear(16, 8, dtype=dtype) self.linear8 = nn.Linear(8, latent_dims, dtype=dtype) def forward(self, x): z = torch.tanh(self.linear1(x)) z = torch.tanh(self.linear2(z)) z = torch.tanh(self.linear3(z)) z = torch.tanh(self.linear4(z)) z = torch.tanh(self.linear5(z)) z = torch.tanh(self.linear6(z)) z = torch.tanh(self.linear7(z)) z = torch.tanh(self.linear8(z)) return z
解码器
class SparseDecoder(nn.Module): def __init__(self, input_shape: int, latent_dims, dtype=torch.float64): super().__init__() self.linear1 = nn.Linear(latent_dims, 8, dtype=dtype) self.linear2 = nn.Linear(8, 16, dtype=dtype) self.linear3 = nn.Linear(16, 32, dtype=dtype) self.linear4 = nn.Linear(32, 64, dtype=dtype) self.linear5 = nn.Linear(64, 128, dtype=dtype) self.linear6 = nn.Linear(128, 256, dtype=dtype) self.linear7 = nn.Linear(256, 512, dtype=dtype) self.linear8 = nn.Linear(512, input_shape, dtype=dtype) def forward(self, x): z = torch.tanh(self.linear1(x)) z = torch.tanh(self.linear2(z)) z = torch.tanh(self.linear3(z)) z = torch.tanh(self.linear4(z)) z = torch.tanh(self.linear5(z)) z = torch.tanh(self.linear6(z)) z = torch.tanh(self.linear7(z)) z = torch.tanh(self.linear8(z)) return z
问题定义
这是激活函数饱和引发的输出边界失真问题:
- tanh的输出范围是[-1,1],而输入是MinMax缩放后落在[0,1]区间的数据,范围不匹配导致极端值(靠近0或1)的映射难度提升。
- 多层tanh堆叠后,神经元极易进入饱和区(当输入绝对值大于~2时,tanh趋近于±1,梯度趋近于0),模型无法有效学习极端值的特征映射,最终输出被"挤压"在中间区间,无法触及0.15以下和0.8以上的范围。
解决建议
- 替换解码器输出层激活函数:将解码器最后一层的tanh换成Sigmoid(输出范围[0,1],完美匹配MinMax缩放后的数据分布),或直接去掉激活函数(若后续需灵活调整输出范围)。编码器的中间层可保留tanh,输出层可根据latent空间需求调整。
- 对齐数据与激活函数范围:将MinMax缩放的范围从[0,1]调整为[-1,1],和tanh的输出范围匹配,避免因范围错位导致的饱和。
- 降低模型复杂度:当前8层全连接层过深,易引发梯度消失和激活饱和。可尝试减少层数(如4-5层),或调整每层神经元数量,简化模型结构。
- 改用非饱和激活函数:将tanh替换为ReLU、GELU等非饱和激活函数,避免梯度消失问题,同时配合Batch Normalization稳定训练过程。
- 加权损失优化:若极端值是业务重点,可采用加权MSE损失,给0.15以下和0.8以上的样本更高权重,引导模型优先拟合这类数据。
- 优化初始化策略:针对tanh使用Xavier初始化,针对ReLU类激活使用He初始化,避免初始权重导致的激活饱和。
内容的提问来源于stack exchange,提问作者Martin Kuhn
相关产品推荐
相关产品推荐

