基于PettingZoo与PyTorch的自定义多智能体环境构建报错求助
问题
我正在为自制游戏构建多智能体模型,打算用PettingZoo将游戏转换为环境,基于PyTorch编写模型。每个智能体包含4种动作:
- Forward:-1到+1的浮点数
- Right:-1到+1的浮点数
- Rotation Rate:-π到+π的浮点数
- Firing:布尔值
观测采用类LiDAR的辐条系统,每个辐条包含0-1比例的距离、检测到的玩家生命值、是否同队,此外还有自身生命值和当前是否可开火。我用Dict类型定义了观测空间与动作空间,但运行封装后的环境校验代码时出现报错。
环境定义代码:
obs_space = Dict({ "distances": Box(low=0, high=1, shape=(8,), dtype=np.float32), "healths": Box(low=0, high=100, shape=(8,), dtype='i'), "alliances": Box(low=-1, high=1, shape=(8,), dtype='i'), "player_health": Discrete(101, start=0), "can_fire": MultiBinary(1), }) self.observation_spaces = {a: obs_space for a in self.agents} action_space = Dict({ "forward": Box(low=-1, high=1, shape=(1,), dtype=np.float32), "right": Box(low=-1, high=1, shape=(1,), dtype=np.float32), "rot_rate": Box(low=-pi, high=pi, shape=(1,), dtype=np.float32), "firing": MultiBinary(1), }) self.action_spaces = {a: action_space for a in self.agents}
运行代码:
from momentum_env.momentum_env import Momentum from torchrl.envs.libs.pettingzoo import PettingZooWrapper env = Momentum() env = PettingZooWrapper(env=env, return_state=False, use_mask=False) check_env_specs(env)
报错核心问题:
TorchRL的PettingZooWrapper对空间类型兼容性要求严格,具体冲突点包括:
- Dict空间中混合Discrete与Box/MultiBinary类型,导致类型转换失败
- Box空间使用整数dtype(
'i'),不符合TorchRL默认的浮点类型期望 - MultiBinary(1)的类型定义存在歧义,与TorchRL的空间处理逻辑不兼容
修复方案
针对上述问题,逐一调整空间定义:
1. 统一观测空间类型,替换Discrete为Box
将player_health的Discrete类型改为Box,适配TorchRL的Dict空间处理逻辑:
# 替换原观测空间中的player_health定义 "player_health": Box(low=0, high=100, shape=(1,), dtype=np.float32),
2. 修正Box空间的整数dtype问题
把healths和alliances的dtype从'i'改为浮点类型;如果alliances仅存在-1/0/1三种取值,也可以用MultiDiscrete更精准定义:
# 浮点类型版本 "healths": Box(low=0, high=100, shape=(8,), dtype=np.float32), "alliances": Box(low=-1, high=1, shape=(8,), dtype=np.float32), # 若alliances只有-1/0/1,推荐用MultiDiscrete # "alliances": MultiDiscrete([3]*8),
3. 替换MultiBinary(1)为Discrete(2)
动作和观测中的布尔值用Discrete(2)定义更清晰,也更适配TorchRL类型系统:
# 观测空间中替换can_fire "can_fire": Discrete(2), # 动作空间中替换firing "firing": Discrete(2),
4. 完整修正后的空间定义
# 观测空间 obs_space = Dict({ "distances": Box(low=0, high=1, shape=(8,), dtype=np.float32), "healths": Box(low=0, high=100, shape=(8,), dtype=np.float32), "alliances": Box(low=-1, high=1, shape=(8,), dtype=np.float32), "player_health": Box(low=0, high=100, shape=(1,), dtype=np.float32), "can_fire": Discrete(2), }) self.observation_spaces = {a: obs_space for a in self.agents} # 动作空间 action_space = Dict({ "forward": Box(low=-1, high=1, shape=(1,), dtype=np.float32), "right": Box(low=-1, high=1, shape=(1,), dtype=np.float32), "rot_rate": Box(low=-np.pi, high=np.pi, shape=(1,), dtype=np.float32), "firing": Discrete(2), }) self.action_spaces = {a: action_space for a in self.agents}
额外注意事项
- 确保环境返回的观测数据类型与空间定义完全匹配,比如
player_health需返回浮点值(如50.0而非50) - 若必须保留Discrete类型,需手动为TorchRL注册空间转换规则,但优先统一为浮点Box类型可大幅简化兼容问题
内容的提问来源于stack exchange,提问作者tensor
相关产品推荐
相关产品推荐

