You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch构建Experience Buffer时遇原地操作错误的解决问询

PyTorch强化学习经验缓冲器梯度错误:原地操作导致的RuntimeError解决

问题背景

从TensorFlow转用PyTorch开发CartPole-v1强化学习模型时,集成经验缓冲器(Replay Buffer)后,训练阶段第二次执行loss.backward()时触发如下错误:

RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation: [torch.FloatTensor [128, 1]], which is output 0 of AsStridedBackward0, is at version 2; expected version 1 instead. Hint: the backtrace further above shows the operation that failed to compute its gradient. The variable in question was changed in there or anywhere later. Good luck!

核心流程:

  • 创建包含网络、优化器、经验缓冲器的Worker类
  • 用collections.deque填充经验缓冲器
  • 训练时仅执行两次采样与反向传播,第二次触发错误

错误底层原因

填充经验缓冲器时,你直接将带计算图的PyTorch张量(action_probs[action]、critic_value)存入了缓冲。这些张量与网络参数的计算图绑定,第一次反向传播后,optimizer.step()会原地更新网络参数(属于inplace操作),导致这些张量关联的计算图节点版本发生变化。第二次反向传播时,PyTorch尝试基于旧版本的计算图节点计算梯度,就会出现版本不匹配的错误。

解决方法

存储经验时必须断开张量与计算图的关联,使用detach()方法生成独立于计算图的张量,确保后续反向传播时,这些张量不会依赖原计算图的参数节点,避免版本冲突。同时修正原代码中损失函数的参数顺序错误,保证训练逻辑正确。

关键修改点

  1. 填充缓冲时处理张量:在将张量存入缓冲前,调用detach()断开计算图,可选.clone()确保张量完全独立
  2. 修正损失函数参数顺序:原代码中ActorLoss的参数传参顺序与定义不匹配,会导致损失计算逻辑偏离预期

完整修正代码

# Load Dependencies
import gym
import numpy as np
import torch as t
import torch.nn as nn
import torch.nn.functional as f
import collections

env = gym.make("CartPole-v1")  # Create the environment
buffer_cap = 100

class Network(nn.Module):
    
    def __init__(self):
        super(Network, self).__init__()

        self.dense1 = nn.Linear(4, 128)
        self.action = nn.Linear(128, 2)
        self.critic = nn.Linear(128, 1)        
        
     
    def forward(self,x):
        x = self.dense1(x)
        x = f.relu(x)
        act = self.action(x)
        act = f.softmax(act, dim = 0)
        crt = self.critic(x)        
        return act, crt 
    
    def ActorLoss(self, log_prob, ret, value):
        # 修正参数逻辑:ret为奖励值,value为critic输出,计算优势函数
        ret = t.tensor(ret, dtype=t.float).unsqueeze(0)
        diff = ret - value
        return -log_prob * diff


#Experience Buffer
class ExpBuffer:
    def __init__(self, capacity):
        self.buffer = collections.deque(maxlen=capacity)

    def append(self, experience):
        self.buffer.append(experience)
        
    def __len__(self):
        return len(self.buffer)

    def sample(self, batch_size):
        indices = np.random.choice(len(self.buffer), batch_size, replace = False) 
        log_probs, values, rewards = zip(*[self.buffer[idx] for idx in indices])
        return zip(log_probs, values, rewards)


class Worker():
    
    def __init__(self):        
        self.net = Network() 
        self.optimizer = t.optim.Adam(self.net.parameters(), lr = 0.02)
        self.ExpBuffer  = ExpBuffer(buffer_cap)         
        self.Experience = collections.namedtuple('Experience', "log_probs value reward")   
        
    def fillExperience(self):
        while len(self.ExpBuffer) < buffer_cap:
            action_probs_hist = []  
            critic_value_hist = []
            rewards_hist = []
            state_next = env.reset()
            done = False    
            while not done:            
                state = t.tensor(state_next, dtype = t.float)
                action_probs, critic_value = self.net.forward(state)
                action = t.multinomial(action_probs,1)
                state_next, reward, done, _ = env.step(int(action))
          
                # 关键修改:detach断开计算图,避免后续参数更新影响
                action_probs_hist.append(action_probs[action].detach().clone())
                critic_value_hist.append(critic_value.detach().clone())
                rewards_hist.append(reward)

            for log_prob, val, rew in zip(action_probs_hist,critic_value_hist,rewards_hist):
                exp = self.Experience(log_prob, val, rew)
                self.ExpBuffer.append(exp)
            
    def train(self):
        for i in range(2):
            loss =  []
            batch = self.ExpBuffer.sample(5)            
                    
            for log_prob, val, ret in batch:
                # 修正参数传入顺序:log_prob, ret(奖励), val(critic输出)
                loss.append(self.net.ActorLoss(log_prob, ret, val))
       
            loss_value = sum(loss)            
            self.optimizer.zero_grad()            
            loss_value.backward()            
            self.optimizer.step()
                
if __name__ == "__main__":
    w = Worker()
    w.fillExperience()
    w.train()            

额外说明

  • detach()会返回一个无梯度信息的张量,原张量的计算图不受影响
  • 若后续需要对存储的张量重新计算梯度,可在训练时将其包装为requires_grad=True的张量,但本场景下历史经验无需关联当前计算图,无需此操作

内容的提问来源于stack exchange,提问作者KO4all

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 09:00:40