Keras转PyTorch后简单测试模型得分差异悬殊,求技术指导
Hey there! Let's dig into why your PyTorch model is underperforming so drastically compared to Keras—there are a few common framework-specific gotchas that are almost certainly causing this huge score gap. Let's break them down and fix them step by step:
1. You're not clearing gradients between batches (the #1 culprit!)
Unlike Keras, which automatically resets gradients after each batch, PyTorch accumulates gradients by default. If you don't explicitly zero them out, the gradients from previous batches keep piling up, leading to completely broken weight updates. This is almost certainly the main reason your model isn't learning.
Fix this by adding optimizer.zero_grad() right before computing the loss for each batch:
for epoch in range(epochs): avg_cost = 0 for i in range(n_batch): x_batch = x_train[i*batch_size:(i+1)*batch_size] y_batch = y_train[i*batch_size:(i+1)*batch_size] # Note: Variable is deprecated in modern PyTorch, tensors handle autograd directly x, y = torch.from_numpy(x_batch).float(), torch.from_numpy(y_batch).float() # Critical: Zero out gradients before each batch optimizer.zero_grad() pred = model(x) loss = loss_fn(pred, y) loss.backward() optimizer.step() avg_cost += loss.item()/n_batch print(epoch, avg_cost)
2. You're dropping the last partial batch + not shuffling data
Your code calculates n_batch = int(x_train.shape[0]/batch_size), which ignores any leftover samples that don't fill a full batch. Keras's fit() automatically handles these partial batches, so your PyTorch model is training on fewer examples than the Keras one. Also, Keras shuffles training data by default—your original code doesn't, which hurts training stability.
Use PyTorch's DataLoader to handle both issues cleanly:
from torch.utils.data import TensorDataset, DataLoader # Wrap your data into a dataset and loader train_dataset = TensorDataset(torch.from_numpy(x_train).float(), torch.from_numpy(y_train).float()) train_loader = DataLoader(train_dataset, batch_size=32, shuffle=True) # Rewrite your training loop to use the loader for epoch in range(epochs): avg_cost = 0 for x, y in train_loader: optimizer.zero_grad() pred = model(x) loss = loss_fn(pred, y) loss.backward() optimizer.step() avg_cost += loss.item() / len(train_loader) print(epoch, avg_cost)
3. Model initialization differences
Keras's Dense layer uses Glorot uniform (Xavier) initialization by default, while PyTorch's Linear layer uses Kaiming uniform (He) initialization for ReLU activations. While this isn't the main issue, aligning initialization can help match Keras's starting behavior:
# Helper function to apply Glorot uniform initialization def init_weights(m): if isinstance(m, torch.nn.Linear): torch.nn.init.xavier_uniform_(m.weight) torch.nn.init.zeros_(m.bias) # Apply initialization to your model model.apply(init_weights)
4. Switch to evaluation mode during inference
Even without Dropout/BatchNorm layers, it's a best practice to set your model to evaluation mode when making predictions. This ensures any training-specific behaviors are disabled, and we can also disable autograd for faster inference:
model.eval() # Switch to evaluation mode with torch.no_grad(): # Turn off autograd to save memory/compute x = torch.from_numpy(x_test).float() pred = model(x).numpy() score(y_test, pred) model.train() # Switch back to training mode if needed later
5. Double-check data consistency
Make sure your training/test data is identical between the Keras and PyTorch code. Verify you're not accidentally normalizing data in one framework but not the other—Keras doesn't apply normalization by default, so as long as you're using the same raw arrays, this should be fine.
After applying these fixes, your PyTorch model's performance should match (or come very close to) your Keras model's 0.9 score. Let me know if you still hit snags!
内容的提问来源于stack exchange,提问作者jl303

