You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

强化学习CartPole模型预测报错:观测形状不匹配问题求助

CartPole-v1模型预测报错:观测形状(2,)不符合要求的(4,)

问题背景

刚入门强化学习,训练CartPole-v1模型后取得500分(远超200分及格线),保存训练文件与模型文件后,测试时调用model.predict出现报错,错误提示观测形状为(2,),不符合Box环境要求的(4,)或(n_env,4)。检查发现env.reset()返回的obs长度为2,疑惑为何不是模型期望的4维,期望解决问题实现CartPole平衡。

测试代码

environment_name = "CartPole-v1"
env = gym.make(environment_name, render_mode='human')

episodes = 5
for episodes in range(1, episodes + 1):
    obs = env.reset()
    done = False
    score = 0

    while not done:
        env.render()

        # this is the key change
        action, _ = model.predict(obs)  # we're now using our model here!
        obs, reward, done, info = env.step(action)

        score += reward
    print("Episode: {} Score: {}".format(episodes, score))
env.close()

报错信息

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
Input In [14], in <cell line: 2>()
      8 env.render()
     10 # this is the key change
---> 11 action, _= model.predict(obs) # we're now using our model here!
     12 obs, reward, done,

ValueError: Error: Unexpected observation shape (2,) for Box environment, please use (4,) or (n_env, 4) for the observation shape.

验证代码及结果

env = gym.make('CartPole-v1')
obs = env.reset()
len(obs)

返回结果:2

问题原因

Gym 0.26及以上版本对env.reset()的返回值做了改动:旧版本仅返回4维的观测数组,新版本返回**(观测值, 额外信息字典)**的元组。你当前拿到的obs是这个元组(长度为2),而非真正的4维观测数据,因此模型报错。

解决方法

方法1:直接解构返回值

修改测试代码中obs = env.reset()这一行,改为:

obs, _ = env.reset()

这样obs会正确获取到4维的观测数组,满足模型输入要求。

方法2:兼容多版本写法

如果需要适配不同Gym版本,可添加判断逻辑:

reset_output = env.reset()
if isinstance(reset_output, tuple):
    obs, _ = reset_output
else:
    obs = reset_output

修改后重新运行验证代码,len(obs)将返回4,此时调用model.predict(obs)即可正常执行。

内容的提问来源于stack exchange,提问作者rj21

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 19:32:36