Keras符号张量转NumPy数组报错:无法转换符号tf.Tensor
问题:Keras自定义损失函数中Tensor转NumPy数组的性能问题
问题背景
使用Python 3.12.3、NumPy 1.26.4、无GPU支持的TensorFlow 2.17.0在私有环境开发,需实现调用C库的Keras自定义损失函数,计划将损失函数参数转为NumPy数组后通过PyBind11传递给C。但在损失函数中执行numpy.asarray(yPred)时触发报错:
NotImplementedError: Cannot convert a symbolic tf.Tensor (data_1:0) to a numpy array. This error may indicate that you're trying to pass a Tensor to a NumPy call, which is not supported.
启用tensorflow.config.run_functions_eagerly(True)可解决报错,但运行时间从11秒骤增至145秒,性能严重下降,寻求高效解决方案。
示例代码
from sklearn.preprocessing import MinMaxScaler from sklearn.metrics import mean_squared_error from keras.layers import Input from keras.layers import Dense from keras import Model import numpy import tensorflow from tensorflow.python.ops import math_ops def pinnLoss(yTrue, yPred): yArr = numpy.asarray(yPred) # Will be used later. squared_difference = math_ops.square(yTrue - yPred) return math_ops.mean(squared_difference, axis=-1) # Note the `axis=-1` x = numpy.asarray([i for i in range(-50,51)]) y = numpy.asarray([i * i for i in x]) x = x.reshape((len(x), 1)) y = y.reshape((len(y), 1)) scale_x = MinMaxScaler() x = scale_x.fit_transform(x) scale_y = MinMaxScaler() y = scale_y.fit_transform(y) inputShape = (1,) inputs = Input(shape=inputShape) tmp = inputs tmp = Dense(10, activation='relu', kernel_initializer='he_uniform')(tmp) tmp = Dense(10, activation='relu', kernel_initializer='he_uniform')(tmp) tmp = Dense(1)(tmp) model = Model(inputs, tmp) model.compile(loss=pinnLoss, optimizer='adam') model.fit(x, y, epochs=500, batch_size=10, verbose=0)
完整报错信息
2024-10-14 11:37:04.117027: I external/local_xla/xla/tsl/cuda/cudart_stub.cc:32] Could not find cuda drivers on your machine, GPU will not be used. 2024-10-14 11:37:04.119994: I external/local_xla/xla/tsl/cuda/cudart_stub.cc:32] Could not find cuda drivers on your machine, GPU will not be used. 2024-10-14 11:37:04.129464: E external/local_xla/xla/stream_executor/cuda/cuda_fft.cc:485] Unable to register cuFFT factory: Attempting to register factory for plugin cuFFT when one has already been registered 2024-10-14 11:37:04.144761: E external/local_xla/xla/stream_executor/cuda/cuda_dnn.cc:8454] Unable to register cuDNN factory: Attempting to register factory for plugin cuDNN when one has already been registered 2024-10-14 11:37:04.149407: E external/local_xla/xla/stream_executor/cuda/cuda_blas.cc:1452] Unable to register cuBLAS factory: Attempting to register factory for plugin cuBLAS when one has already been registered 2024-10-14 11:37:04.975719: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT Traceback (most recent call last): File "/home/bamer/work/dev/optimizer/src/minimalworking2.py", line 31, in <module> model.fit(x, y, epochs=500, batch_size=10, verbose=0) File "/home/bamer/ownpython/lib/python3.12/site-packages/keras/src/utils/traceback_utils.py", line 122, in error_handler raise e.with_traceback(filtered_tb) from None File "/home/bamer/work/dev/optimizer/src/minimalworking2.py", line 11, in pinnLoss yArr = numpy.asarray(yPred) # Will be used later. ^^^^^^^^^^^^^^^^^^^^ NotImplementedError: Cannot convert a symbolic tf.Tensor (data_1:0) to a numpy array. This error may indicate that you're trying to pass a Tensor to a NumPy call, which is not supported.
高效解决方案
方法1:用tf.py_function封装C++调用逻辑
该方法在TensorFlow图模式下执行,同时支持调用Python函数(进而调用C++库),避免全Eager模式的性能损耗。
实现代码
from sklearn.preprocessing import MinMaxScaler from keras.layers import Input from keras.layers import Dense from keras import Model import numpy import tensorflow as tf from tensorflow.python.ops import math_ops # 替换为你的PyBind11调用C++库的逻辑 def cpp_loss_calculation(y_true_np, y_pred_np): # 这里写实际调用C++库的代码 squared_diff = (y_true_np - y_pred_np) ** 2 return numpy.mean(squared_diff, axis=-1) def pinnLoss(yTrue, yPred): # 用tf.py_function包装Python函数,指定输入输出类型 loss = tf.py_function( func=cpp_loss_calculation, inp=[yTrue, yPred], Tout=tf.float32 ) # 保持张量形状,避免后续模型训练报错 loss.set_shape(yTrue.get_shape()) return loss # 后续模型构建与训练代码不变 x = numpy.asarray([i for i in range(-50,51)]) y = numpy.asarray([i * i for i in x]) x = x.reshape((len(x), 1)) y = y.reshape((len(y), 1)) scale_x = MinMaxScaler() x = scale_x.fit_transform(x) scale_y = MinMaxScaler() y = scale_y.fit_transform(y) inputShape = (1,) inputs = Input(shape=inputShape) tmp = inputs tmp = Dense(10, activation='relu', kernel_initializer='he_uniform')(tmp) tmp = Dense(10, activation='relu', kernel_initializer='he_uniform')(tmp) tmp = Dense(1)(tmp) model = Model(inputs, tmp) model.compile(loss=pinnLoss, optimizer='adam') model.fit(x, y, epochs=500, batch_size=10, verbose=0)
方法2:实现TensorFlow自定义Op(极致性能)
若追求最高性能,可将C++逻辑封装为TensorFlow自定义Op,完全融入TensorFlow图计算,性能与原生Op一致。
步骤
- 用C++编写符合TensorFlow Op规范的代码,包含Op注册逻辑
- 将代码编译为动态链接库(.so文件)
- 在Python中加载该库,直接在损失函数中调用自定义Op
此方法门槛较高,但适合计算密集型的损失函数场景。
方法3:结合TensorFlow数据管道批量处理
若C++库支持批量处理,可使用tf.data管道提前将数据转为NumPy数组处理后喂给模型,适合离线预处理场景,但实时损失计算仍推荐前两种方法。
内容的提问来源于stack exchange,提问作者Balázs Bämer
相关产品推荐
相关产品推荐

