You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CartPole-v1神经网络训练遇TypeError问题求助

解决CartPole-v1环境中神经网络的TypeError问题

问题代码

env = gym.make("CartPole-v1",render_mode="rgb_array")
obs = env.reset()

n_inputs = env.observation_space.shape[0]

model = tf.keras.Sequential([
    tf.keras.layers.Dense(5, activation="relu"),
    tf.keras.layers.Dense(1, activation="sigmoid"),
])

def play_one_step(env, obs, model, loss_function):
    
    with tf.GradientTape() as tape:
        
        left_probability = model(obs[np.newaxis])
        action = (tf.random.uniform([1, 1]) > left_probability)
        y_target = tf.constant([[1.]]) - tf.cast(action, tf.float32)
        loss = tf.reduce_mean(loss_function(y_target, left_probability))
        
    gradients = tape.gradient(loss, model.trainable_variables)
    obs, reward, done, truncated, info = env.step(int(action))
    
    return obs, reward, done, truncated, gradients

报错信息

Cell In [158], line 30, in play_episodes(env, n_episodes, n_max_steps, model, loss_function)
     26 obs = env.reset()
     28 for step in range(n_max_steps):
---> 30     obs, reward, done, truncated, gradients = play_one_step(env, obs, model, loss_function)
     31     current_rewards.append(reward)
     32     current_gradients.append(gradients)

Cell In [158], line 7, in play_one_step(env, obs, model, loss_function)
      3 def play_one_step(env, obs, model, loss_function):
      5     with tf.GradientTape() as tape:
----> 7         left_probability = model(obs[np.newaxis])
      8         action = (tf.random.uniform([1, 1]) > left_probability)
      9         y_target = tf.constant([[1.]]) - tf.cast(action, tf.float32)

TypeError: tuple indices must be integers or slices, not NoneType

问题根源

新版本Gym库中,当指定render_mode参数时,env.reset()会返回**(观测值, 信息字典)**的元组,而非单独的观测值。代码中直接将obs赋值为env.reset()的结果,导致obs是一个元组,后续执行obs[np.newaxis]时触发错误——元组的索引只能是整数或切片,不能是None(np.newaxis本质是None)。

之前尝试的形状调整、转Tensor操作无效,都是因为操作对象是元组而非真正的观测数组。

修复方案

1. 正确提取观测值

修改env.reset()的赋值语句,从返回的元组中提取观测值:

obs, _ = env.reset()  # 忽略不需要的info字典

2. 确保模型输入维度正确

Keras模型期望输入是**(batch_size, 特征数)**的2D张量,处理单个观测时,需要扩展batch维度:

# 方式1:用numpy扩展维度
left_probability = model(obs[np.newaxis, :])
# 方式2:用TensorFlow扩展维度
left_probability = model(tf.expand_dims(obs, 0))

修复后的完整代码片段

env = gym.make("CartPole-v1", render_mode="rgb_array")
obs, _ = env.reset()  # 修正:提取观测值

n_inputs = env.observation_space.shape[0]

model = tf.keras.Sequential([
    tf.keras.layers.Dense(5, activation="relu"),
    tf.keras.layers.Dense(1, activation="sigmoid"),
])

def play_one_step(env, obs, model, loss_function):
    
    with tf.GradientTape() as tape:
        # 修正:正确扩展输入维度
        left_probability = model(obs[np.newaxis, :])
        action = (tf.random.uniform([1, 1]) > left_probability)
        y_target = tf.constant([[1.]]) - tf.cast(action, tf.float32)
        loss = tf.reduce_mean(loss_function(y_target, left_probability))
        
    gradients = tape.gradient(loss, model.trainable_variables)
    obs, reward, done, truncated, info = env.step(int(action))
    
    return obs, reward, done, truncated, gradients

内容的提问来源于stack exchange,提问作者Ravi Sharma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 04:08:28