You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于PettingZoo与PyTorch的自定义多智能体环境构建报错求助

问题

我正在为自制游戏构建多智能体模型,打算用PettingZoo将游戏转换为环境,基于PyTorch编写模型。每个智能体包含4种动作:

  • Forward:-1到+1的浮点数
  • Right:-1到+1的浮点数
  • Rotation Rate:-π到+π的浮点数
  • Firing:布尔值

观测采用类LiDAR的辐条系统,每个辐条包含0-1比例的距离、检测到的玩家生命值、是否同队,此外还有自身生命值和当前是否可开火。我用Dict类型定义了观测空间与动作空间,但运行封装后的环境校验代码时出现报错。

环境定义代码:

obs_space = Dict({
    "distances": Box(low=0, high=1, shape=(8,), dtype=np.float32),
    "healths": Box(low=0, high=100, shape=(8,), dtype='i'),
    "alliances": Box(low=-1, high=1, shape=(8,), dtype='i'),
    "player_health": Discrete(101, start=0),
    "can_fire": MultiBinary(1),
})

self.observation_spaces = {a: obs_space for a in self.agents}

action_space = Dict({
    "forward": Box(low=-1, high=1, shape=(1,), dtype=np.float32),
    "right": Box(low=-1, high=1, shape=(1,), dtype=np.float32),
    "rot_rate": Box(low=-pi, high=pi, shape=(1,), dtype=np.float32),
    "firing": MultiBinary(1),
})

self.action_spaces = {a: action_space for a in self.agents}

运行代码:

from momentum_env.momentum_env import Momentum
from torchrl.envs.libs.pettingzoo import PettingZooWrapper

env = Momentum()
env = PettingZooWrapper(env=env, return_state=False, use_mask=False)

check_env_specs(env)

报错核心问题:
TorchRL的PettingZooWrapper对空间类型兼容性要求严格,具体冲突点包括:

  1. Dict空间中混合Discrete与Box/MultiBinary类型,导致类型转换失败
  2. Box空间使用整数dtype('i'),不符合TorchRL默认的浮点类型期望
  3. MultiBinary(1)的类型定义存在歧义,与TorchRL的空间处理逻辑不兼容

修复方案

针对上述问题,逐一调整空间定义:

1. 统一观测空间类型,替换Discrete为Box

将player_health的Discrete类型改为Box,适配TorchRL的Dict空间处理逻辑:

# 替换原观测空间中的player_health定义
"player_health": Box(low=0, high=100, shape=(1,), dtype=np.float32),

2. 修正Box空间的整数dtype问题

把healths和alliances的dtype从'i'改为浮点类型;如果alliances仅存在-1/0/1三种取值,也可以用MultiDiscrete更精准定义:

# 浮点类型版本
"healths": Box(low=0, high=100, shape=(8,), dtype=np.float32),
"alliances": Box(low=-1, high=1, shape=(8,), dtype=np.float32),

# 若alliances只有-1/0/1,推荐用MultiDiscrete
# "alliances": MultiDiscrete([3]*8),

3. 替换MultiBinary(1)为Discrete(2)

动作和观测中的布尔值用Discrete(2)定义更清晰,也更适配TorchRL类型系统:

# 观测空间中替换can_fire
"can_fire": Discrete(2),

# 动作空间中替换firing
"firing": Discrete(2),

4. 完整修正后的空间定义

# 观测空间
obs_space = Dict({
    "distances": Box(low=0, high=1, shape=(8,), dtype=np.float32),
    "healths": Box(low=0, high=100, shape=(8,), dtype=np.float32),
    "alliances": Box(low=-1, high=1, shape=(8,), dtype=np.float32),
    "player_health": Box(low=0, high=100, shape=(1,), dtype=np.float32),
    "can_fire": Discrete(2),
})

self.observation_spaces = {a: obs_space for a in self.agents}

# 动作空间
action_space = Dict({
    "forward": Box(low=-1, high=1, shape=(1,), dtype=np.float32),
    "right": Box(low=-1, high=1, shape=(1,), dtype=np.float32),
    "rot_rate": Box(low=-np.pi, high=np.pi, shape=(1,), dtype=np.float32),
    "firing": Discrete(2),
})

self.action_spaces = {a: action_space for a in self.agents}

额外注意事项

  • 确保环境返回的观测数据类型与空间定义完全匹配,比如player_health需返回浮点值(如50.0而非50)
  • 若必须保留Discrete类型,需手动为TorchRL注册空间转换规则,但优先统一为浮点Box类型可大幅简化兼容问题

内容的提问来源于stack exchange,提问作者tensor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 21:08:22