强化学习CartPole模型预测报错:观测形状不匹配问题求助
CartPole-v1模型预测报错:观测形状(2,)不符合要求的(4,)
问题背景
刚入门强化学习,训练CartPole-v1模型后取得500分(远超200分及格线),保存训练文件与模型文件后,测试时调用model.predict出现报错,错误提示观测形状为(2,),不符合Box环境要求的(4,)或(n_env,4)。检查发现env.reset()返回的obs长度为2,疑惑为何不是模型期望的4维,期望解决问题实现CartPole平衡。
测试代码
environment_name = "CartPole-v1" env = gym.make(environment_name, render_mode='human') episodes = 5 for episodes in range(1, episodes + 1): obs = env.reset() done = False score = 0 while not done: env.render() # this is the key change action, _ = model.predict(obs) # we're now using our model here! obs, reward, done, info = env.step(action) score += reward print("Episode: {} Score: {}".format(episodes, score)) env.close()
报错信息
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) Input In [14], in <cell line: 2>() 8 env.render() 10 # this is the key change ---> 11 action, _= model.predict(obs) # we're now using our model here! 12 obs, reward, done, ValueError: Error: Unexpected observation shape (2,) for Box environment, please use (4,) or (n_env, 4) for the observation shape.
验证代码及结果
env = gym.make('CartPole-v1') obs = env.reset() len(obs)
返回结果:2
问题原因
Gym 0.26及以上版本对env.reset()的返回值做了改动:旧版本仅返回4维的观测数组,新版本返回**(观测值, 额外信息字典)**的元组。你当前拿到的obs是这个元组(长度为2),而非真正的4维观测数据,因此模型报错。
解决方法
方法1:直接解构返回值
修改测试代码中obs = env.reset()这一行,改为:
obs, _ = env.reset()
这样obs会正确获取到4维的观测数组,满足模型输入要求。
方法2:兼容多版本写法
如果需要适配不同Gym版本,可添加判断逻辑:
reset_output = env.reset() if isinstance(reset_output, tuple): obs, _ = reset_output else: obs = reset_output
修改后重新运行验证代码,len(obs)将返回4,此时调用model.predict(obs)即可正常执行。
内容的提问来源于stack exchange,提问作者rj21
相关产品推荐
相关产品推荐

