OpenAI Gym step函数解包5变量仍报错:ValueError(预期5值仅得4个)
问题:Gym环境step方法解包报错(期望5个值,实际得到4个)
按照官方文档编写代码后仍触发过时错误,代码如下:
done = True env.reset() for step in range(5000): action = env.action_space.sample() obs, reward, terminated, truncated, info = env.step(action) done = terminated or truncated if done: env.reset() env.close()
报错信息:
ValueError Traceback (most recent call last) /UsersProjects/GAME_AI_RL/Super_Mario/main.ipynb Cell 8 line 5 3 for step in range(5000): 4 action = env.action_space.sample() ----> 5 obs, reward, terminated, truncated, info = env.step(action) 6 done = terminated or truncated 8 if done: File ~/miniforge3/envs/mlp/lib/python3.8/site-packages/nes_py/wrappers/joypad_space.py:74, in JoypadSpace.step(self, action) 59 """ 60 Take a step using the given action. 61 (...) 71 72 """ 73 # take the step and record the output ---> 74 return self.env.step(self._action_map[action]) File ~/miniforge3/envs/mlp/lib/python3.8/site-packages/gym/wrappers/time_limit.py:50, in TimeLimit.step(self, action) 39 def step(self, action): 40 """Steps through the environment and if the number of steps elapsed exceeds <code>max_episode_steps</code> then truncate. 41 42 Args: (...) ... ---> 50 observation, reward, terminated, truncated, info = self.env.step(action) 51 self._elapsed_steps += 1 53 if self._elapsed_steps >= self._max_episode_steps: ValueError: not enough values to unpack (expected 5, got 4)
运行环境说明:MACOS M1,基于虚拟环境。最初使用4个变量解包,查询得知接口已更新为5变量后修改,但仍报错;已尝试重启内核、重新安装库、更换Python版本。
解决方案
核心问题是你使用的nes_py库(对应Super Mario环境)未适配Gym新API——Gym新API的step方法返回5个值(obs, reward, terminated, truncated, info),但该库底层仍返回旧API的4个值(obs, reward, done, info)。
可选方案:
方案1:回退到Gym旧版本
安装Gym 0.25.x及以前的版本,使用旧API的4变量解包逻辑:done = True env.reset() for step in range(5000): action = env.action_space.sample() obs, reward, done, info = env.step(action) if done: env.reset() env.close()方案2:切换到Gymnasium+适配新API的环境包
使用Gymnasium(Gym的官方继任者),同时安装支持新API的gym-super-mario-bros最新版本,代码保持你当前的5变量解包逻辑即可。方案3:手动封装Wrapper转换API格式
自己写一个Wrapper将旧API的返回值转换为新API格式,无需修改环境依赖:import gym from gym import Wrapper class OldToNewAPIStepWrapper(Wrapper): def step(self, action): obs, reward, done, info = self.env.step(action) # 旧API的done对应新API的terminated,truncated默认设为False(如需区分可自行扩展) terminated = done truncated = False return obs, reward, terminated, truncated, info # 创建环境后应用该Wrapper env = gym.make('SuperMarioBros-v0') env = OldToNewAPIStepWrapper(env)
内容的提问来源于stack exchange,提问作者Rishit Chugh
相关产品推荐
相关产品推荐

