You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow中获取指定迭代MSE值的问题:原因及解决方法

问题:TensorFlow中无法获取迭代时MSE的实际数值

我在使用TensorFlow做梯度下降训练时,想要在特定的epoch和batch组合时打印MSE的实际数值,但直接打印mse得到的是Tensor对象(示例输出:Epoch 0 Batch_Index 0 MSE: Tensor("mse_2:0", shape=(), dtype=float32))。我理解这是因为MSE依赖tf.placeholder节点,但我已经在sess.run(training_op, feed_dict={X: X_batch, y: y_batch})里传入了数据,本以为此时可以获取MSE的值,结果用mse.eval()的时候报错:

InvalidArgumentError: You must feed a value for placeholder tensor 'X_2' with dtype float and shape [?,9]...

为什么会出现这种情况?该怎么修改代码才能输出指定迭代时的MSE值?

附相关代码片段:

import numpy as np
from sklearn.datasets import fetch_california_housing
housing = fetch_california_housing()
m, n = housing.data.shape
housing_data_plus_bias = np.c_[np.ones((m, 1)), housing.data] # ADD COLUMN OF 1s for BIAS!
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaled_housing_data = scaler.fit_transform(housing.data)
scaled_housing_data_plus_bias = np.c_[np.ones((m, 1)), scaled_housing_data]
X = tf.placeholder(tf.float32, shape=(None, n + 1), name="X")
y = tf.placeholder(tf.float32, shape=(None, 1), name="y")
theta = tf.Variable(tf.random_uniform([n + 1, 1], -1.0, 1.0, seed=42), name="theta")
y_pred = tf.matmul(X, theta, name="predictions")
error = y_pred - y
mse = tf.reduce_mean(tf.square(error), name="mse")
optimizer = tf.train.GradientDescentOptimizer(learning_rate=learning_rate)
training_op = optimizer.minimize(mse)
init = tf.global_variables_initializer()
n_epochs = 100
batch_size = 100
n_batches = int(np.ceil(m / batch_size))
learning_rate = 0.01

def fetch_batch(epoch, batch_index, batch_size):
    np.random.seed(epoch * n_batches + batch_index) # not shown in the book
    indices = np.random.randint(m, size=batch_size) # not shown
    X_batch = scaled_housing_data_plus_bias[indices] # not shown
    y_batch = housing.target.reshape(-1, 1)[indices] # not shown
    return X_batch, y_batch

with tf.Session() as sess:
    sess.run(init)
    for epoch in range(n_epochs):
        for batch_index in range(n_batches):
            X_batch, y_batch = fetch_batch(epoch, batch_index, batch_size)
            sess.run(training_op, feed_dict={X: X_batch, y: y_batch})
            if (epoch % 50 == 0 and batch_index % 100 == 0):
                print("Epoch", epoch, "Batch_Index", batch_index, "MSE:", mse)
    best_theta = theta.eval()

原因分析

其实这是TensorFlow惰性执行机制导致的,给你拆解下:

  • 你定义的mse只是计算图里的一个节点,它本身不存储任何数值,每次要获取它的结果,都必须给它依赖的所有占位符(也就是X和y)喂入数据。
  • 你之前调用sess.run(training_op)的时候,确实传入了feed_dict,但这只是让TensorFlow执行了training_op对应的计算(也就是更新参数theta),并没有同时计算mse的值,更不会把这次喂的数据缓存下来供后续使用。mse.eval()本质上等价于sess.run(mse),但这时候你没传feed_dict,TensorFlow找不到X和y的数据,自然就报错了。

解决方法

这里有两种常用的方式,你可以根据需求选:

方法一:同一次sess.run()中同时执行训练和计算MSE(推荐)

这种方式效率更高,因为不需要重复喂数据:
修改循环内的代码,把训练和MSE计算合并到一次sess.run()调用中:

with tf.Session() as sess:
    sess.run(init)
    for epoch in range(n_epochs):
        for batch_index in range(n_batches):
            X_batch, y_batch = fetch_batch(epoch, batch_index, batch_size)
            if (epoch % 50 == 0 and batch_index % 100 == 0):
                # 同时执行训练操作和计算MSE,返回对应的结果
                _, current_mse = sess.run([training_op, mse], feed_dict={X: X_batch, y: y_batch})
                print("Epoch", epoch, "Batch_Index", batch_index, "MSE:", current_mse)
            else:
                # 只执行训练操作
                sess.run(training_op, feed_dict={X: X_batch, y: y_batch})
    best_theta = theta.eval()

sess.run()可以接收一个节点列表,一次性执行多个计算,返回对应顺序的结果。这里_用来接收training_op的返回值(它没有实际意义,因为训练操作只是更新参数),current_mse就是当前batch的MSE实际数值。

方法二:单独计算MSE并传入feed_dict

如果调试时想单独查看MSE,也可以在需要打印的时候单独调用sess.run(mse),但必须再次传入feed_dict:

with tf.Session() as sess:
    sess.run(init)
    for epoch in range(n_epochs):
        for batch_index in range(n_batches):
            X_batch, y_batch = fetch_batch(epoch, batch_index, batch_size)
            sess.run(training_op, feed_dict={X: X_batch, y: y_batch})
            if (epoch % 50 == 0 and batch_index % 100 == 0):
                # 单独计算MSE,记得传入feed_dict
                current_mse = sess.run(mse, feed_dict={X: X_batch, y: y_batch})
                # 或者用 mse.eval(feed_dict={X: X_batch, y: y_batch}) 效果一样
                print("Epoch", epoch, "Batch_Index", batch_index, "MSE:", current_mse)
    best_theta = theta.eval()

这种方式会多一次MSE的计算过程,效率比方法一低一点,但逻辑更直观,适合调试阶段使用。


内容的提问来源于stack exchange,提问作者sebtac

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:40:45