You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于PyTorch Sequential搭建无Gym依赖的实时游戏强化学习模型的相关疑问

基于PyTorch Sequential搭建无Gym依赖的实时游戏强化学习模型的相关疑问

嗨,看起来你已经在手动搭建实时游戏RL模型的路上走了不少路啦,先给你点个赞!我来帮你梳理下代码里的疑问和可以优化的地方:

关于优化器和模型的关联问题

你担心优化器没和模型关联?其实完全不用担心!在PyTorch里,当你执行optimizer = optim.Adam(model.parameters(), lr=1E-2)的时候,已经把model里所有可训练参数的引用传给优化器了。优化器内部会跟踪这些参数的梯度,当你调用optimizer.step()时,就会自动更新model里的参数——这是PyTorch内置的机制,不用手动做额外的关联操作~

代码里需要修正的关键问题

你的整体思路是对的,但还有不少细节会导致代码运行报错或者训练无效:

  • 缺失关键导入:代码里用到了Categorical来采样动作,但没导入这个类,必须加上:

    from torch.distributions import Categorical
    
  • 临时break导致训练无法进行:你在while not dead循环里加了break # Temp,这会让每个episode只运行1步就退出,根本没法让agent和游戏持续交互,记得把这个break删掉。

  • 模型维度不匹配:你最后一层全连接层的输入维度3721计算错误,必须重新计算卷积后的特征图尺寸:
    预处理后的输入是3×256×256,经过两次卷积+池化后:

    1. 第一次Conv2d(3,256,5) → 尺寸为256×252×252,MaxPool2d(2,2) → 256×126×126
    2. 第二次Conv2d(256,512,5) → 尺寸为512×122×122,MaxPool2d(2,2) → 512×61×61
      Flatten之后的总维度是512×61×61=1905152,所以全连接层要改成:
    nn.Linear(512*61*61, 512),
    
  • 奖励逻辑和折扣回报的问题:

    • 当前rewards.append(1 if not dead else 0)的逻辑有问题:dead是在get()里检测的,当检测到死亡时,这一步的动作已经导致失败,奖励应该设为负数(比如-10)而不是1,这样模型才能学会避免导致死亡的动作。
    • 你注释掉的compute_discounted_rewards函数是强化学习的核心(用来计算折扣回报,让模型更关注长期收益),建议排查它的报错原因(比如当rewards只有1个元素时,标准差为0的情况,可以加个判断处理),然后替换掉当前的discounted_rewards = rewards,不然模型训练效率会极低。
  • 死亡检测和输入参数问题:

    • get(raw=True)调用了不存在的参数,你的get()函数没有定义raw参数,要改成get()。
    • 基于RGB均值的死亡检测逻辑不够可靠,比如游戏场景中可能出现类似的颜色分布但并未死亡,建议换成更精准的判断(比如检测死亡画面中固定位置的像素颜色,或者用简单的文字识别匹配死亡提示)。
  • 冗余代码:while循环结束后的dead = False是冗余的,因为下一个episode开头已经重新设置了dead = False,可以删掉。

修正后的核心代码片段示例

这里给你改了关键部分的代码参考:

from PIL import ImageGrab, Image
import numpy as np
import torch
from torch import nn
from torchvision import transforms
import torch.optim as optim
from torch.distributions import Categorical  # 新增导入
import keyboard
import pyautogui

_bbox=(800, 200, 1200, 900) # (L,T,R,B)
dead = False


def get():
    global dead
    screenshot = ImageGrab.grab(bbox=_bbox)

    # 优化死亡检测逻辑(示例,你可以根据实际游戏画面调整)
    screen_array = np.array(screenshot)
    r_mean = np.mean(screen_array[:,:,0])
    g_mean = np.mean(screen_array[:,:,1])
    b_mean = np.mean(screen_array[:,:,2])
    if r_mean >= 190 and g_mean <50 and b_mean <15:
        dead = True

    return screenshot.copy()


def compute_discounted_rewards(rewards, gamma=0.99):
    discounted_rewards = []
    R = 0
    for r in reversed(rewards):
        R = r + gamma * R
        discounted_rewards.insert(0, R)
    discounted_rewards = torch.tensor(discounted_rewards)
    # 处理单元素的情况,避免标准差为0
    if len(discounted_rewards) ==1:
        return discounted_rewards
    discounted_rewards = (discounted_rewards - discounted_rewards.mean()) / (discounted_rewards.std() + 1e-5)
    return discounted_rewards

# 定义预处理
preprocess = transforms.Compose([
    transforms.Resize(256),
    transforms.CenterCrop(256),
    transforms.ToTensor(),
    transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
])

# 修正后的模型结构
model = nn.Sequential(
          nn.Conv2d(3,256,5),
          nn.ReLU(),
          nn.MaxPool2d(2, 2),
          nn.Conv2d(256,512,5),
          nn.ReLU(),  # 把ReLU移到MaxPool后面,更合理
          nn.MaxPool2d(2, 2),
          nn.Flatten(),
          nn.Linear(512*61*61, 512),
          nn.LeakyReLU(),
          nn.Linear(512, 512),
          nn.LeakyReLU(),
          nn.Linear(512, 3),
          nn.Softmax(dim=1)
        )

optimizer = optim.Adam(model.parameters(), lr=1E-4)  # 把学习率调低一点,1E-2太大容易震荡
ep_reward = list()

for j in range(10):  # 增加episode数量,至少跑几十次才能看到效果
    dead = False
    i = 0
    log_probs = list()
    rewards = list()
    while not dead:
        screenshot = get()
        inputs = preprocess(screenshot).unsqueeze(0)  # 增加batch维度,模型需要batch输入
        outputs = model(inputs)
        m = Categorical(outputs)
        action = m.sample()

        log_probs.append(m.log_prob(action))
        # 调整奖励:活着就给1分,死亡给-10分
        rewards.append(1 if not dead else -10)
        
        # 这里根据action执行游戏操作,比如:
        # if action.item() ==1:
        #     pyautogui.press('left')
        # elif action.item() ==2:
        #     pyautogui.press('right')
        # else:
        #     pass
        
        i += 1
    
    ep_reward.append(sum(rewards))
    discounted_rewards = compute_discounted_rewards(rewards)
    policy_loss = []
    for log_prob, Gt in zip(log_probs, discounted_rewards):
        policy_loss.append(-log_prob * Gt)
    optimizer.zero_grad()
    policy_loss = torch.cat(policy_loss).sum()
    policy_loss.backward()
    optimizer.step()
    print(f"Episode {j+1}, Total Reward: {sum(rewards)}")

整体来说,你的思路是完全可行的,只要把这些细节问题修正,模型就能正常运行并开始学习啦!

备注:内容来源于stack exchange,提问作者Controller816

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 15:28:08