RLlib多智能体环境观测空间配置错误排查求助
问题排查与解决:RLlib多智能体环境观测空间不匹配错误
错误根源分析
你遇到的ValueError核心原因是观测空间定义与实际返回的观测结构不匹配:
- 当前把
observation_space定义为包含两个智能体容量的全局字典,但RLlib的MultiAgentEnv要求观测空间是以智能体ID为键,每个智能体独立观测空间为值的字典。 step/reset返回的观测应该是{智能体ID: 该智能体的观测数据}结构,而不是直接返回全局容量字典——此时每个智能体的观测被错误当成整个全局字典,而RLlib期望每个智能体的观测是单个张量/向量。- 代码中还存在小bug:
rewards字典的键混用了字符串("1")和整数(1),会导致后续匹配错误。
修复步骤
- 重新定义每个智能体的观测空间:每个智能体需要观测自身和对方的容量,因此每个智能体的观测空间是形状为(2,)的Box,范围[0,100]。
- 修正全局观测空间结构:用
Dict包裹每个智能体的独立观测空间,键对应智能体ID。 - 修正观测返回格式:在
reset和step中,为每个智能体生成包含自身+对方容量的观测数组。 - 修正done的返回格式:MultiAgentEnv要求done是包含
__all__键的字典,表示全局是否结束。 - 统一rewards的键类型:全部使用字符串类型的智能体ID。
修正后的环境代码
from ray.rllib.env.multi_agent_env import MultiAgentEnv from gymnasium.spaces import Discrete, Box, Dict class CapacityEnv(MultiAgentEnv): def __init__(self): # 每个智能体的动作空间都是离散2选1 self.action_space = Discrete(2) # 0: 不传输, 1: 传输 # 每个智能体的观测空间:[自身容量, 对方容量],范围0-100 self.observation_space = Dict({ "1": Box(low=0, high=101, shape=(2,), dtype=int), "2": Box(low=0, high=101, shape=(2,), dtype=int) }) self.node_capacity = {"1": 100, "2": 100} def step(self, action_dict): node_choice_1 = action_dict["1"] node_choice_2 = action_dict["2"] rewards = {"1": 0, "2": 0} # 双方选择相同动作,惩罚 if node_choice_1 == node_choice_2: rewards = {"1": -10, "2": -10} # 节点2选择传输,节点1不传输 elif node_choice_1 == 0 and node_choice_2 == 1: if self.node_capacity["2"] >= self.node_capacity["1"]: rewards = {"1": 10, "2": 10} else: rewards = {"1": -10, "2": -10} self.node_capacity["2"] -= 5 # 节点1选择传输,节点2不传输 elif node_choice_1 == 1 and node_choice_2 == 0: if self.node_capacity["1"] >= self.node_capacity["2"]: rewards = {"1": 10, "2": 10} else: rewards = {"1": -10, "2": -10} self.node_capacity["1"] -= 5 # 生成每个智能体的观测:[自身容量, 对方容量] observations = { "1": [self.node_capacity["1"], self.node_capacity["2"]], "2": [self.node_capacity["2"], self.node_capacity["1"]] } # 判断全局是否结束 done_flag = self.node_capacity["1"] == 0 or self.node_capacity["2"] == 0 done = {"__all__": done_flag} return observations, rewards, done, False, {} def reset(self, *, seed=None, options=None): self.node_capacity = {"1": 100, "2": 100} # 生成初始观测 observations = { "1": [self.node_capacity["1"], self.node_capacity["2"]], "2": [self.node_capacity["2"], self.node_capacity["1"]] } return observations, {}
修正后的训练代码
from ray.rllib.algorithms.dqn import DQNConfig from ray.tune.logger.logger import pretty_print # 配置多智能体训练:为每个智能体使用相同的DQN策略 config = DQNConfig().environment(CapacityEnv).training(gamma=0.9, lr=0.001, train_batch_size=512) # 显式指定多智能体策略映射(如果不指定,RLlib会自动为每个智能体创建独立策略) config.multi_agent( policies={"shared_policy": (None, config.environment_config.observation_space, config.environment_config.action_space, {})}, policy_mapping_fn=lambda agent_id, *args, **kwargs: "shared_policy" ) agent = config.build() for i in range(5): result = agent.train() print(f"训练迭代 {i+1}:") print(pretty_print(result))
关键说明
- 观测空间现在与实际返回的观测结构完全匹配:每个智能体的观测是长度为2的数组,对应观测空间定义的(2,)形状的Box。
- 多智能体策略配置中,这里使用了共享策略(两个智能体共用同一个DQN策略),你也可以改为为每个智能体创建独立策略,只需要调整
policy_mapping_fn即可。 - 修正了
done的返回格式,符合RLlib对MultiAgentEnv的要求。
内容的提问来源于stack exchange,提问作者Laura
相关产品推荐
相关产品推荐

