You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

加载PyTorch权重后PPO模型性能骤降,与训练效果差距显著

PPO模型加载后表现异常问题

我基于一套PPO强化学习教程完成CartPole-v1实验,训练阶段一切正常,模型能够稳定保持杆直立。但加载训练好的模型后,其表现完全等同于随机初始化的模型。已确认检查点文件的时间戳是最新的,说明训练后的权重已经覆盖了旧文件。

我仅有的修改是在ActorNetwork和CriticNetwork类中,为self.checkpoint_file添加了.pth后缀,修改后的路径定义如下:

self.checkpoint_file = os.path.join(chkpt_dir, 'actor_torch_ppo.pth')
#...
self.checkpoint_file = os.path.join(chkpt_dir, 'critic_torch_ppo.pth')

我的推理代码inference.py内容如下:

import gym
import numpy as np
from ppo_torch import Agent
import time
import sys

env = gym.make('CartPole-v1', render_mode = "human")

batch_size = 16
n_epochs = 5
alpha = 0.0003

agent = Agent(n_actions=env.action_space.n, batch_size=batch_size, 
                    alpha=alpha, n_epochs=n_epochs, 
                    input_dims=env.observation_space.shape)

agent.load_models()
agent.critic.eval()
agent.actor.eval()

while True:
    observation, _ = env.reset()
    done = False
    while not done:
        start = time.time()
        action, prob, val = agent.choose_action(observation)
        observation_, reward, done, info, _ = env.step(action)

        env.render()
        end = time.time()
        print(f'{(1/(end-start)):.2f}', end="\r", flush=True)

内容的提问来源于stack exchange,提问作者explodingfilms101

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 02:00:09