torch.autograd.grad计算LSTM输出对时间导数返回None求助
问题:LSTM与PINN结合时无法计算输出对时间的梯度
我有一个LSTM模型,输入3个温度数据序列,输出下一个序列:
input => [array([0.20408163, 0.40816327, 0.6122449 ]), array([0.40816327, 0.6122449 , 0.81632653])] output=> [tensor(0.81632653, dtype=torch.float64), tensor(0.91667510, dtype=torch.float64)]
我希望将该LSTM模型与基于牛顿冷却定律的物理信息神经网络(PINN)结合,思路是用LSTM预测温度,再计算预测温度对时间的导数,将物理定律融入损失函数。但尝试计算LSTM输出对时间t的梯度时,torch.autograd.grad返回None,不确定是否正确使用了torch.autograd。
以下是简化代码:
import torch import torch.nn as nn def create_lstm_model(input_size, hidden_size, num_layers, output_size): class LSTMModel(nn.Module): def __init__(self, input_size, hidden_size, num_layers, output_size): super(LSTMModel, self).__init__() self.hidden_size = hidden_size self.num_layers = num_layers self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True) self.fc = nn.Linear(hidden_size, output_size) def forward(self, x): h0 = torch.zeros(self.num_layers, x.size(0), self.hidden_size).to(x.device) c0 = torch.zeros(self.num_layers, x.size(0), self.hidden_size).to(x.device) out, _ = self.lstm(x, (h0, c0)) out = self.fc(out[:, -1, :]) return out return LSTMModel(input_size, hidden_size, num_layers, output_size) def physics_loss_autograd(outputs, time_step): """ Compute the physics-informed loss using autograd to get dT/dt. """ # Compute dT/dt using autograd dT_dt = torch.autograd.grad(outputs, time_step, grad_outputs=torch.ones_like(outputs), create_graph=True)[0] # Newton's law of cooling: dT/dt = -k(T - T_ambient) residual = dT_dt + k * (outputs - T_ambient) # Physics loss is the L2 norm of the residual physics_loss = torch.mean(residual**2) return physics_loss t = torch.arange(0,100) input_size = 1 hidden_size = 64 num_layers = 1 output_size = 1 # Create an instance of the LSTMModel using the function model = create_lstm_model(input_size, hidden_size, num_layers, output_size) criterion = nn.MSELoss() optimizer = torch.optim.Adam(model.parameters(), lr=0.001) num_epochs = 100 for epoch in range(num_epochs): model.train() total_loss = 0 for inputs, targets in train_loader: inputs, targets = inputs.float(), targets.float() # Convert to float # print(inputs.shape) optimizer.zero_grad() outputs = model(inputs) data_loss = criterion(outputs, targets) # DO SOMETHING LIKE phys_loss = physics_loss_autograd(outputs, t) loss = data_loss + phys_loss loss.backward() optimizer.step() total_loss += loss.item() if (epoch+1) % 20 == 0: print(f'Epoch {epoch+1}/{num_epochs}, Loss: {total_loss/len(train_loader)}')
我尝试过将时间步作为额外特征输入LSTM,但仍无法计算导数。请求指导如何计算LSTM输出的时间导数,解决grad返回None的问题。
解决方案
问题根源
torch.autograd.grad返回None的核心原因是:当前的LSTM输出与时间变量t之间没有计算图连接。你的代码中t只是一个独立的数组,没有作为模型输入的一部分参与前向传播,因此PyTorch无法追踪输出对t的梯度。
修正步骤
1. 将时间步作为模型输入的一部分
需要让每个输入序列对应明确的时间戳,将时间特征与温度序列一起输入LSTM。例如,每个输入样本可以是[温度值, 对应时间步],调整输入处理逻辑:
# 假设train_loader中的每个样本包含温度序列、对应时间序列和目标值 for temp_seq, time_seq, targets in train_loader: # 将温度序列和时间序列拼接,输入维度变为2(温度+时间) inputs = torch.cat([temp_seq.unsqueeze(-1), time_seq.unsqueeze(-1)], dim=-1).float() targets = targets.float() optimizer.zero_grad() outputs = model(inputs) # ...后续损失计算与反向传播
同时调整模型的输入维度:
input_size = 2 # 温度特征 + 时间特征
2. 确保时间变量可求导
必须将时间序列转换为可求导的张量,设置requires_grad=True让PyTorch追踪其梯度:
# 在数据加载时标记时间序列可求导 time_seq = time_seq.float().requires_grad_(True)
3. 修正物理损失计算函数
调整梯度计算逻辑,确保保留计算图以支持后续反向传播:
def physics_loss_autograd(outputs, time_seq): # 计算dT/dt,设置retain_graph=True保留计算图 dT_dt = torch.autograd.grad( outputs, time_seq, grad_outputs=torch.ones_like(outputs), create_graph=True, retain_graph=True )[0] # 定义牛顿冷却定律的参数(可设为常数或可学习参数) k = torch.tensor(0.1, requires_grad=False) T_ambient = torch.tensor(25.0, requires_grad=False) residual = dT_dt + k * (outputs - T_ambient) physics_loss = torch.mean(residual**2) return physics_loss
4. 确认模型前向传播逻辑
确保模型能正确处理多维度输入:
def create_lstm_model(input_size, hidden_size, num_layers, output_size): class LSTMModel(nn.Module): def __init__(self, input_size, hidden_size, num_layers, output_size): super(LSTMModel, self).__init__() self.hidden_size = hidden_size self.num_layers = num_layers self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True) self.fc = nn.Linear(hidden_size, output_size) def forward(self, x): # x形状:[batch_size, seq_len, input_size] h0 = torch.zeros(self.num_layers, x.size(0), self.hidden_size).to(x.device) c0 = torch.zeros(self.num_layers, x.size(0), self.hidden_size).to(x.device) out, _ = self.lstm(x, (h0, c0)) # 取最后一个时间步输出预测下一个温度 out = self.fc(out[:, -1, :]) return out return LSTMModel(input_size, hidden_size, num_layers, output_size)
关键注意事项
- 必须保证时间变量参与模型前向传播,形成完整计算图,才能计算输出对时间的梯度。
- 用
requires_grad=True标记时间张量,确保PyTorch追踪其梯度信息。 torch.autograd.grad中设置create_graph=True和retain_graph=True,保证后续损失反向传播正常进行。
内容的提问来源于stack exchange,提问作者Abdul Rehman
相关产品推荐
相关产品推荐

