TensorFlow Keras图执行错误求助:遵循教程训练仍报错
问题描述
我知道已经有类似问题被提出,但我严格按照教程操作还是出现错误,查遍了所有可能的解决方案仍不知道怎么修正。我初步判断问题和形状或标签有关,但找不到解决办法,求帮忙。
我的代码
import numpy as np import matplotlib.pyplot as plt import tensorflow as tf from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense, Input, Activation from tensorflow.keras.datasets import boston_housing from tensorflow.keras import layers SEED_VALUE = 65 # 固定随机种子保证结果可复现 np.random.seed(SEED_VALUE) tf.random.set_seed(SEED_VALUE) # 加载波士顿房价数据集 (X_train, y_train), (X_test, y_test ) = boston_housing.load_data() print(X_train.shape) print("\n") print("输入特征样例: ", X_train[0]) print("\n") print("标签样例: ", y_train[0]) boston_features = { 'Average Number of Rooms': 5, } # 提取单特征(平均房间数) X_train_1d = X_train[:, boston_features['Average Number of Rooms']] print(X_train_1d.shape) X_test_1d = X_test[:, boston_features['Average Number of Rooms']] # 绘制特征与标签的散点图 plt.figure(figsize=(15,5)) plt.xlabel('平均房间数') plt.ylabel('房价中位数 [$K]') plt.grid("on") plt.scatter(X_train_1d[:], y_train, color='green', alpha=0.5); # 构建单神经元模型 model = Sequential() model.add(Dense(units=1, input_shape=(1,))) # 打印模型结构 model.summary() # 编译模型 model.compile(optimizer=tf.keras.optimizers.RMSprop(learning_rate=.005), loss='mse') # 训练模型 history = model.fit(X_train_1d, y_train, batch_size=16, epochs=101, validation_split=0.3)
运行错误信息
Epoch 1/101 2023-01-21 12:03:45.701983: I tensorflow/core/grappler/optimizers/custom_graph_optimizer_registry.cc:114] Plugin optimizer for device_type GPU is enabled. 2023-01-21 12:03:45.757923: W tensorflow/core/framework/op_kernel.cc:1830] OP_REQUIRES failed at xla_ops.cc:418 : NOT_FOUND: could not find registered platform with id: 0x281ae11b0 2023-01-21 12:03:45.757952: W tensorflow/core/framework/op_kernel.cc:1830] OP_REQUIRES failed at xla_ops.cc:418 : NOT_FOUND: could not find registered platform with id: 0x281ae11b0 --------------------------------------------------------------------------- NotFoundError Traceback (most recent call last) Cell In[28], line 1 ----> 1 history = model.fit(X_train_1d, y_train, batch_size=16, epochs=101, validation_split=0.3) File ~/miniconda3/envs/tensorflow/lib/python3.10/site-packages/keras/utils/traceback_utils.py:70, in filter_traceback.<locals>.error_handler(*args, **kwargs) 67 filtered_tb = _process_traceback_frames(e.__traceback__) 68 # To get the full stack trace, call: 69 # `tf.debugging.disable_traceback_filtering()` ---> 70 raise e.with_traceback(filtered_tb) from None 71 finally: 72 del filtered_tb File ~/miniconda3/envs/tensorflow/lib/python3.10/site-packages/tensorflow/python/eager/execute.py:52, in quick_execute(op_name, num_outputs, inputs, attrs, ctx, name) 50 try: 51 ctx.ensure_initialized() ---> 52 tensors = pywrap_tfe.TFE_Py_Execute(ctx._handle, device_name, op_name, 53 inputs, attrs, num_outputs) 54 except core._NotOkStatusException as e: 55 if name is not None: NotFoundError: Graph execution error:
解决方案
这个错误和数据形状、标签无关,是TensorFlow的GPU适配问题,可通过以下方式解决:
- 强制用CPU运行:在代码开头添加以下代码,禁用GPU,让TensorFlow使用CPU计算:
import os os.environ["CUDA_VISIBLE_DEVICES"] = "-1"
检查版本兼容性:当前TensorFlow版本和系统的CUDA、cuDNN版本不匹配,导致XLA加速模块找不到GPU平台。对照TensorFlow官方版本兼容表,安装对应版本的CUDA和cuDNN。
关闭XLA优化:添加配置禁用XLA,避免相关错误:
tf.config.optimizer.set_jit(False)
- 升级TensorFlow:部分旧版本存在GPU平台注册bug,升级到最新稳定版可修复:
pip install --upgrade tensorflow
内容的提问来源于stack exchange,提问作者Andres Beregovich
相关产品推荐
相关产品推荐

