You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ray Tune多智能体环境中同一模型参数重复采样问题求助

问题排查与解决方案

1. 修复代码语法错误

你的代码中pol_1的配置元组存在语法错误:act_space_a后面缺少逗号,会导致Python解析失败,先补上这个逗号:

"pol_1": (
    None,
    obs_space_a,
    act_space_a,  # 补上该逗号
    {
        "model": {
            "custom_model": "model_a",
            "custom_model_config": {
                "hidden_layer_size": 64,
                "num_hidden_layers": 2,
                "activation": "leaky_relu",
            },
        },
    },
),

2. 超参数配置方式错误(核心问题)

直接在policies字典中嵌入tune.choice会导致Ray Tune与RLlib的参数解析逻辑冲突:当你把config.to_dict()传给Tuner的param_space时,Tune会遍历所有嵌套的tune变量,但RLlib在初始化多智能体策略时,会对策略配置做内部复制与处理,导致同一个超参数被重复采样。

正确的配置方式

先定义基础策略配置(用默认值填充待调参的字段),再在param_space中明确覆盖pol_2的超参数:

if __name__ == "__main__":
    ray.init()

    # 基础策略配置:pol_2用默认值占位,后续在param_space中覆盖
    policies = {
        "pol_1": (
            None,
            obs_space_a,
            act_space_a,
            {
                "model": {
                    "custom_model": "model_a",
                    "custom_model_config": {
                        "hidden_layer_size": 64,
                        "num_hidden_layers": 2,
                        "activation": "leaky_relu",
                    },
                },
            },
        ),
        "pol_2": (
            None,
            obs_space_b,
            act_space_a,
            {
                "model": {
                    "custom_model": "attacker_model",
                    "custom_model_config": {
                        "hidden_layer_size": 64,  # 默认占位值
                        "num_hidden_layers": 2,
                    },
                },
            },
        ),
    }

    config = (
        AlgorithmConfig()
        .environment("concurrent_env", env_config={"num_agents": 2})
        .training(train_batch_size=1024, lr=1e-3, gamma=0.99)
        .framework("torch")
        .rollouts(num_rollout_workers=1, rollout_fragment_length="auto")
        .multi_agent(
            policies=policies,
            policy_mapping_fn=(
                lambda agent_id, episode, worker, **kw: f"pol_{agent_id}"
            ),
            policies_to_train=["pol_1", "pol_2"],
        )
    )

    # 生成param_space并覆盖pol_2的超参数
    param_space = config.to_dict()
    param_space["multi_agent"]["policies"]["pol_2"]["model"]["custom_model_config"]["hidden_layer_size"] = tune.choice([64, 128, 256, 512, 1024])

    results = tune.Tuner(
        "PPO",
        param_space=param_space,
        run_config=air.RunConfig(
            stop={"training_iteration": 10000},
            callbacks=[AttackerRewardCallback()],
        ),
        tune_config=tune.TuneConfig(num_samples=1),
    ).fit()

3. 验证效果

修改后,Tune只会对pol_2的hidden_layer_size进行一次采样,且自定义模型的forward方法会正确使用该采样值。


内容的提问来源于stack exchange,提问作者Pat396

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 16:10:42