You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于TensorFlow解决OpenAI Gym BipedalWalker任务报错咨询

问题修复方案

1 报错根因说明

你对报错原因的猜测是错误的,IndexError: invalid index to scalar variable的核心问题是你选择的强化学习算法不匹配任务属性:DQNAgent仅支持离散动作空间任务,而BipedalWalker-v3的动作空间是4维连续值,DQN内部会尝试对输出的动作值做索引取离散动作,自然触发索引报错。

2 核心修改点

  • 更换为支持连续动作空间的DDPG算法,适配BipedalWalker的动作属性
  • 将模型输出层的激活函数更换为tanh,其输出天然在[-1,1]区间,直接满足输出范围要求

3 修改后的可运行完整代码

import gym
from tensorflow.keras.models import Sequential, Model
from tensorflow.keras.layers import Dense, Flatten, Activation, Input, Concatenate
from tensorflow.keras.optimizers import Adam
from rl.agents import DDPGAgent
from rl.memory import SequentialMemory
from rl.random import OrnsteinUhlenbeckProcess

env = gym.make("BipedalWalker-v3")
states = env.observation_space.shape[0]
actions = env.action_space.shape[0]

# 构建演员模型(负责输出动作)
def build_actor(states, actions):
    model = Sequential()
    model.add(Flatten(input_shape=(1, states)))
    model.add(Dense(24, activation='relu'))
    model.add(Dense(24, activation='relu'))
    model.add(Dense(actions))
    model.add(Activation('tanh')) # tanh激活直接限制所有输出在[-1,1]区间
    return model

# 构建评论家模型(负责评价动作价值)
def build_critic(states, actions):
    action_input = Input(shape=(actions,), name='action_input')
    observation_input = Input(shape=(1, states), name='observation_input')
    flattened_observation = Flatten()(observation_input)
    x = Concatenate()([action_input, flattened_observation])
    x = Dense(32, activation='relu')(x)
    x = Dense(32, activation='relu')(x)
    x = Dense(1)(x)
    return Model(inputs=[action_input, observation_input], outputs=x)

actor = build_actor(states, actions)
critic = build_critic(states, actions)
action_input = Input(shape=(actions,))
memory = SequentialMemory(limit=100000, window_length=1)
# 加入随机探索过程提升训练效果
random_process = OrnsteinUhlenbeckProcess(size=actions, theta=.15, mu=0., sigma=.3)

# 初始化DDPG智能体
agent = DDPGAgent(
    nb_actions=actions,
    actor=actor,
    critic=critic,
    critic_action_input=action_input,
    memory=memory,
    nb_steps_warmup_critic=100,
    nb_steps_warmup_actor=100,
    random_process=random_process,
    gamma=.99,
    target_model_update=1e-3
)

agent.compile(Adam(learning_rate=1e-3, clipnorm=1.), metrics=['mae'])
# 训练模型,可根据效果调整训练步数
agent.fit(env, nb_steps=500000, visualize=False, verbose=1, nb_max_episode_steps=200)
# 测试训练效果
agent.test(env, nb_episodes=15, visualize=True, nb_max_episode_steps=200)
# 保存权重
agent.save_weights('ddpg_bipedal_weights.h5f', overwrite=True)

4 可选方案说明

如果你坚持要使用DQN算法,可以手动将4维连续动作空间离散化,比如每个维度拆成-1、0、1三档,总共有81个离散动作后传入DQN训练,但该方案的训练效果远不如直接使用连续动作专属算法。

内容的提问来源于stack exchange,提问作者Malte Rothkamm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 02:48:03