You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow中Adam优化器Eager/非Eager更新差异致RL收敛变慢的原因与解决

问题

我用TensorFlow和Gymnasium做Atari游戏强化学习开发时发现:用tf.function的非Eager执行(延迟执行)比Eager执行的模型收敛需要更多迭代次数,尽管非Eager整体速度更快。排查后发现是优化器的动量和速度变量更新逻辑存在差异,下面是最小复现示例,仅5次训练迭代后,两种模式下的优化器变量就会出现不同:

import tensorflow as tf
import numpy as np
import gymnasium as gym

@tf.function
def discounted_cumulative_sums(
    rewards: tf.Tensor,
    gamma: float,
    standardize: bool = True) -> tf.Tensor:
    """Compute expected returns per timestep."""

    n = tf.shape(rewards)[0]
    returns = tf.TensorArray(dtype=tf.float32, size=n)

    # Start from the end of `rewards` and accumulate reward sums
    # into the `returns` array
    rewards = tf.cast(rewards[::-1], dtype=tf.float32)
    discounted_sum = tf.constant(0.0)
    discounted_sum_shape = discounted_sum.shape
    for i in tf.range(n):
        reward = rewards[i]
        discounted_sum = reward + gamma * discounted_sum
        discounted_sum.set_shape(discounted_sum_shape)
        returns = returns.write(i, discounted_sum)
    returns = returns.stack()[::-1]

    if standardize:
        returns = ((returns - tf.math.reduce_mean(returns)) /
                   tf.math.reduce_std(returns))

    return returns

def make_model(input_shape,output_nodes) : 
    tf.keras.utils.set_random_seed(1956) #set seed so it always returns identical model    
    input_layer = tf.keras.layers.Input(input_shape)
    x = tf.keras.layers.Dense(64,activation = 'relu')(input_layer)
    x = tf.keras.layers.Dense(64,activation = 'relu')(x)
    output_layer = tf.keras.layers.Dense(output_nodes,activation = 'linear')(x)
    model = tf.keras.Model(input_layer,output_layer)
    return model

@tf.function
def train_model(old_states, A, model, optimizer) : 
    for i in tf.range(5) : 
        with tf.GradientTape() as tape : 
            loss = tf.reduce_mean(tf.squeeze(model(old_states),1) * A)

        grads = tape.gradient(loss,model.trainable_variables)
        optimizer.apply_gradients(zip(grads,model.trainable_variables))

    return optimizer.variables

# create data
td_error = tf.random.normal(shape=(100,))
old_states = tf.random.normal(shape=(100,8))

# create models and optimizers
eager_model = make_model((8,),1)
non_eager_model = make_model((8,),1)
eager_optimizer = tf.keras.optimizers.Adam(learning_rate = 0.0005)
non_eager_optimizer = tf.keras.optimizers.Adam(learning_rate = 0.0005)

tf.config.run_functions_eagerly(True) # EAGER EXECUTION
A = discounted_cumulative_sums(td_error,.99)
eager_opt_variables = train_model(old_states,A,eager_model,eager_optimizer)

tf.config.run_functions_eagerly(False) # NON-EAGER EXECUTION
A = discounted_cumulative_sums(td_error,.99) # if this is commented out, there will be no differences between optimizer variables
non_eager_opt_variables = train_model(old_states,A,non_eager_model,non_eager_optimizer)

# Show differences between optimizer variables after training
for x in range(22) : 
    print(eager_opt_variables[x].numpy()==non_eager_opt_variables[x].numpy())

运行代码后可见,eager_opt_variables和non_eager_opt_variables多数情况下不相等。差异虽小,但会随迭代累积或放大影响收敛速度。有意思的是,只要注释掉调用discounted_cumulative_sums生成A的代码(比如直接设A = td_error),两种模式下的优化器变量就完全一致。

我的疑问是:

  1. 这个现象的原因是什么?
  2. 怎么让非Eager执行的结果和Eager执行保持一致?

原因分析

核心问题出在浮点数计算精度差异以及tf.function图执行与Eager执行的计算路径细微区别:

  • discounted_cumulative_sums函数包含循环和TensorArray操作,Eager模式下是逐次执行浮点数运算;而图执行(非Eager)会对计算图做优化(比如算子融合、精度对齐微调),导致生成的A张量在浮点数末位出现微小差异。
  • Adam优化器的动量(m)和速度(v)变量依赖梯度累积,初始的微小差异会在多次迭代(示例中的5次循环)中被放大,最终导致优化器变量出现可观测的不同。
  • 当直接用td_error作为A时,张量没有经过复杂计算,两种模式下的A完全一致,因此后续优化器更新也完全同步。

解决方案

要让非Eager执行和Eager执行结果一致,可以从以下几点入手:

  1. 固定计算图的精度行为
    在tf.function装饰器中添加experimental_relax_shapes=False,同时设置全局浮点数精度策略,避免自动精度转换导致的差异:

    tf.keras.mixed_precision.set_global_policy('float32')
    
    @tf.function(experimental_relax_shapes=False)
    def discounted_cumulative_sums(...):
        # 函数内容不变
    
  2. 统一A的生成与训练的执行模式
    把A的生成和训练逻辑放在同一个tf.function里,确保图执行时A的计算和后续训练在同一个图中,消除跨模式的精度差异:

    @tf.function
    def train_with_A(td_error, old_states, model, optimizer, gamma=0.99):
        A = discounted_cumulative_sums(td_error, gamma)
        for i in tf.range(5):
            with tf.GradientTape() as tape:
                loss = tf.reduce_mean(tf.squeeze(model(old_states), 1) * A)
            grads = tape.gradient(loss, model.trainable_variables)
            optimizer.apply_gradients(zip(grads, model.trainable_variables))
        return optimizer.variables
    
  3. 禁用图执行的部分优化(可选)
    如果上述方法无效,可以禁用图执行的算子融合优化,但这会损失部分性能,建议仅在需要严格对齐结果时使用:

    tf.config.optimizer.set_experimental_options({'disable_meta_optimizer': True})
    

内容的提问来源于stack exchange,提问作者Mas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 07:15:16