You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch与Keras等价LSTM模型验证及结果差异排查

TensorFlow(Keras)与PyTorch LSTM模型等价性验证及结果差异排查

一、模型结构等价性验证:完全不等价

两个模型的核心结构存在多处关键差异,是结果不一致的核心原因:

TensorFlow(Keras)模型核心结构

model = Sequential()
model.add(LSTM(5, input_shape=(4, 1000)))  # 单层LSTM,隐藏单元数5,仅返回最后时间步输出
model.add(Dense(1, activation='tanh'))     # 输出层激活为tanh
model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])
  • LSTM层:1层,隐藏单元数5,输入序列长度4、特征数1000
  • 输出层:Dense输出1,激活函数tanh(输出范围[-1,1])
  • 损失函数:binary_crossentropy(该损失期望输出为[0,1]区间的概率值,此处与tanh激活不匹配,本身存在逻辑问题)

PyTorch模型核心结构(与TensorFlow的关键差异)

class LSTM1(nn.Module):
    def __init__(self, num_classes, input_size, hidden_size, num_layers, seq_length):
        super().__init__()
        self.lstm = nn.LSTM(input_size=input_size, hidden_size=hidden_size,
                          num_layers=num_layers, batch_first=True)  # 5层LSTM,隐藏单元数1
        self.fc = nn.Linear(self.hidden_size, num_classes)
        self.sigmoid = nn.Sigmoid()
    
    def forward(self,x):
        h_0 = torch.zeros(self.num_layers, x.size(0), self.hidden_size)
        c_0 = torch.zeros(self.num_layers, x.size(0), self.hidden_size)
        output, (hn, cn) = self.lstm(x, (h_0, c_0))
        hn = hn.view(-1, self.hidden_size)
        out = self.sigmoid(hn)
        out = self.fc(out)
        out = self.sigmoid(out)  # 两次Sigmoid激活
        return out

# 初始化参数
num_layers = 5  # 与TensorFlow的1层完全相反
hidden_size = 1 # 与TensorFlow的5完全相反
  1. LSTM核心参数颠倒:TensorFlow是1层LSTM、5个隐藏单元;PyTorch是5层LSTM、1个隐藏单元,结构完全不同
  2. 激活函数不匹配:TensorFlow输出层用tanh,PyTorch对LSTM输出和最终输出都用Sigmoid,输出范围和特性完全不同
  3. 输出层逻辑冗余:PyTorch中对LSTM隐藏状态先做Sigmoid再过全连接层,属于多余操作,进一步偏离TensorFlow结构

二、训练过程的额外差异(加剧结果不一致)

除了模型结构,训练流程也存在多处错误:

  1. TensorFlow训练的自身问题

    • 输出层用tanh激活,但损失函数是binary_crossentropy:该损失要求输出为[0,1]的概率,而tanh输出是[-1,1],会导致损失计算异常,且准确率指标的判定逻辑(默认以0为阈值,而非0.5)不符合二分类常规逻辑
  2. PyTorch训练的错误操作

    • 标签处理逻辑混乱:每次循环都将y_train_tensors转为LongTensor再转回float,并重复reshape,会干扰梯度计算
    • 多余的输出截断:outputs = outputs[-20:]——输入本身是20个样本,此操作无意义,若后续样本量变化会直接报错
    • 准确率计算逻辑:将标签转为LongTensor后与预测值比较,虽结果数值上可能一致,但类型转换冗余且易引发问题
  3. 初始化与优化细节差异

    • 模型参数初始化策略:Keras与PyTorch的默认初始化方法不同(如Keras LSTM用Glorot均匀初始化,PyTorch用Xavier均匀初始化),会导致初始损失不同,但这是次要因素

三、修正建议(让两个模型等价)

要让PyTorch模型与TensorFlow模型对齐,需做以下修改:

修正后的PyTorch模型

class LSTM1(nn.Module):
    def __init__(self, input_size, hidden_size, num_classes):
        super(LSTM1, self).__init__()
        # 对齐TensorFlow:1层LSTM,hidden_size=5,batch_first=True匹配输入形状(20,4,1000)
        self.lstm = nn.LSTM(input_size=input_size, hidden_size=hidden_size,
                          num_layers=1, batch_first=True)
        self.fc = nn.Linear(hidden_size, num_classes)
        self.tanh = nn.Tanh()  # 对齐TensorFlow的tanh激活
    
    def forward(self,x):
        # 初始隐藏状态与细胞状态全0,对齐Keras默认行为
        h_0 = torch.zeros(1, x.size(0), self.lstm.hidden_size, device=x.device)
        c_0 = torch.zeros(1, x.size(0), self.lstm.hidden_size, device=x.device)
        # 对齐Keras的return_sequences=False,取最后时间步的隐藏状态
        _, (hn, _) = self.lstm(x, (h_0, c_0))
        hn = hn.squeeze(0)  # 形状从(1, batch_size, 5)转为(batch_size,5)
        out = self.fc(hn)
        out = self.tanh(out)
        return out

# 初始化参数对齐TensorFlow
input_size = 1000
hidden_size = 5  # 对应Keras的LSTM(5)
num_classes = 1

model = LSTM1(input_size, hidden_size, num_classes)

修正后的PyTorch训练流程

# 标签处理:直接转为float并调整形状,无需重复转换
y_train_tensors = torch.Tensor(y_train).unsqueeze(1)  # 形状转为(20,1),对齐模型输出

criterion = torch.nn.BCELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.001)  # Adam默认学习率0.001,对齐Keras

for epoch in range(100):
    outputs = model(X_train_tensors_final)
    optimizer.zero_grad()
    
    # 对齐TensorFlow的损失计算:将tanh输出转为[0,1]区间以适配BCELoss
    outputs_normalized = (outputs + 1) / 2
    loss = criterion(outputs_normalized, y_train_tensors)
    
    loss.backward()
    optimizer.step()
    
    # 对齐TensorFlow的准确率逻辑:tanh输出>0时视为1
    preds = (outputs > 0).float()
    train_acc = torch.sum(preds == y_train_tensors) / len(y_train_tensors)
    
    print(f"Epoch: {epoch}, loss: {loss.item():.5f}")
    print(f'Accuracy: {train_acc.item():.4f}')

同时修正TensorFlow模型的自身问题

原TensorFlow模型中tanh激活搭配binary_crossentropy是错误的,建议改为:

model.add(Dense(1, activation='sigmoid'))  # 用Sigmoid输出[0,1]概率,匹配binary_crossentropy

内容的提问来源于stack exchange,提问作者abdullatif

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 17:25:31