You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

本地运行TensorFlow Actor-Critic代码遇OperatorNotAllowedInGraphError报错

TensorFlow Actor-Critic代码本地运行时Graph模式循环报错

将TensorFlow官方Actor-Critic教程的代码从Colab迁移到本地PyCharm运行,即使直接运行原版代码也出现错误,核心问题是循环遍历符号化tf.Tensor时触发报错。

报错代码与信息

触发错误的循环代码:

for t in tf.range(max_steps):
    state = tf.expand_dims(state, 0)
    softmax_action, value = model(state)
    action = tf.random.categorical(softmax_action, 1)[0,0]

报错信息:

File "", line 89, in run_episode
tensorflow.python.framework.errors_impl.OperatorNotAllowedInGraphError:
Iterating over a symbolic tf.Tensor is not allowed: AutoGraph did
convert this function. This might indicate you are trying to use an
unsupported feature.

即使注释循环改为仅运行一次(设置t=0),get_returns()函数内的循环仍会触发相同错误:

File "", line 124, in get_returns
tensorflow.python.framework.errors_impl.OperatorNotAllowedInGraphError:
Iterating over a symbolic tf.Tensor is not allowed: AutoGraph did
convert this function. This might indicate you are trying to use an
unsupported feature.

环境信息

  • 设备:搭载M1芯片的macOS
  • TensorFlow版本:tensorflow-macos2.11.0、tensorflow-metal0.7.1

精简代码片段

def run_episode(state, model, max_steps):
    action_probs = tf.TensorArray(dtype=tf.float32, size=0, dynamic_size=True)
    values = tf.TensorArray(dtype=tf.float32, size=0, dynamic_size=True)
    rewards = tf.TensorArray(dtype=tf.int32, size=0, dynamic_size=True)
    initial_state_shape = state.shape

    for t in tf.range(max_steps):
        state = tf.expand_dims(state, 0)
        softmax_action, value = model(state)
        action = tf.random.categorical(softmax_action, 1)[0,0]
        # ... 后续代码省略...

@tf.function
def train_step(init_state, model, gamma, max_steps):

    with tf.GradientTape() as tape:
        action_probs, values, rewards = run_episode(init_state, model, max_steps)
        returns = get_returns(rewards, gamma, True)

        loss = custom_loss(action_probs, values, returns)

        grads = tape.gradient(loss, model.trainable_variables)
        optimizer.apply_gradients(zip(grads, model.trainable_variables))

        episode_reward = tf.math.reduce_sum(rewards)

        return episode_reward


t = tqdm.trange(num_episodes)
for i in t:
    state, info = env.reset()  # 测试时可注释
    state = tf.constant(state, dtype=tf.float32)


    reward = train_step(state, model, gamma, max_steps) 

内容的提问来源于stack exchange,提问作者brian_ds

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 08:42:26