You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow每100步打印训练准确率,修改代码后无效求助

我帮你排查出问题所在啦!你修改后没生效,大概率是两个容易踩的坑:一是训练模式下没正确计算准确率张量,二是没把日志钩子正确传给EstimatorSpec。咱们一步步来解决:

第一步:修正准确率计算与钩子配置

先把你的cnn_model_fn训练分支代码改成下面这样,关键改动我标了注释:

def cnn_model_fn(features, labels, mode):
    # 保持原有的网络层定义(输入层、卷积层、池化层等)不变
    # ...

    # 先定义预测结果,这部分可以放到所有模式共享的位置
    logits = tf.layers.dense(inputs=flatten, units=10)
    predictions = {
        "classes": tf.argmax(input=logits, axis=1),
        "probabilities": tf.nn.softmax(logits, name="softmax_tensor")
    }

    # 计算损失(所有模式共享)
    loss = tf.losses.sparse_softmax_cross_entropy(labels=labels, logits=logits)

    if mode == tf.estimator.ModeKeys.TRAIN:
        optimizer = tf.train.GradientDescentOptimizer(learning_rate=0.001)
        train_op = optimizer.minimize(
            loss=loss,
            global_step=tf.train.get_global_step())
        
        # 重点1:计算当前batch的训练准确率(用即时计算而非metrics的累计值)
        train_accuracy = tf.reduce_mean(
            tf.cast(tf.equal(predictions["classes"], labels), tf.float32))
        
        # 定义日志钩子,指定每100步打印准确率
        logging_hook = tf.train.LoggingTensorHook(
            {"训练准确率": train_accuracy},
            every_n_iter=100)
        
        # 重点2:必须把钩子传入EstimatorSpec的hooks参数,不然不会生效!
        return tf.estimator.EstimatorSpec(
            mode=mode,
            loss=loss,
            train_op=train_op,
            hooks=[logging_hook])  # 这里一定要加!

    # 以下PREDICT和EVAL模式的代码保持原教程逻辑即可
    if mode == tf.estimator.ModeKeys.PREDICT:
        return tf.estimator.EstimatorSpec(mode=mode, predictions=predictions)

    eval_metric_ops = {
        "accuracy": tf.metrics.accuracy(labels=labels, predictions=predictions["classes"])
    }
    return tf.estimator.EstimatorSpec(
        mode=mode, loss=loss, eval_metric_ops=eval_metric_ops)
第二步:关键细节说明
  • 如果你用tf.metrics.accuracy,它返回的是(更新操作, 累计准确率)的元组,所以要打印的话得取第二个元素accuracy[1];而上面代码里用的是当前batch的即时准确率,更适合训练过程中的实时监控。
  • 最容易漏掉的就是把logging_hook放到EstimatorSpec的hooks列表里——这是钩子能被训练流程调用的关键,你之前的代码应该就是没加这一步,导致钩子根本没触发。
测试效果

修改完启动训练后,你会看到每100步输出类似这样的日志:

INFO:tensorflow:训练准确率 = 0.89
INFO:tensorflow:global_step/sec: 112.34

这样就成功实现每100步记录训练准确率的需求啦!

内容的提问来源于stack exchange,提问作者Caitlin Wen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 11:13:35