R语言Keras混合密度网络自定义损失函数报错求助
解决R Keras混合密度网络自定义损失函数的TypeError错误
问题重现
训练混合密度网络时触发如下错误:
Epoch 1/25 2023-10-23 11:55:50.657605: I tensorflow/core/grappler/optimizers/custom_graph_optimizer_registry.cc:114] Plugin optimizer for device_type GPU is enabled. 750/750 [==============================] - 3s 4ms/step - loss: 75.5268 Error in py_call_impl(callable, call_args$unnamed, call_args$named) : TypeError: in user code: File "/Users/ryanbmac/.virtualenvs/r-tensorflow/lib/python3.9/site-packages/keras/src/engine/training.py", line 1972, in test_function * return step_function(self, iterator) File "/Library/Frameworks/R.framework/Versions/4.2-arm64/Resources/library/reticulate/python/rpytools/call.py", line 16, in python_function * raise error File "/Users/ryanbmac/.virtualenvs/r-tensorflow/lib/python3.9/site-packages/keras/src/backend.py", line 3613, in reshape return tf.reshape(x, shape) TypeError: Failed to convert elements of (None, 1) to Tensor. Consider casting elements to a supported type. See https://www.tensorflow.org/api_docs/python/tf/dtypes for supported TF dtypes.
用户核心代码如下:
K <- 5 # the number of Gaussian components input_shape <- ncol(X_train) # input shape # Define the shared neural network model <- keras_model_sequential() model %>% layer_dense(units = 64, activation = "relu", input_shape = input_shape) %>% layer_dense(units = 64, activation = "relu") %>% layer_dense(units = K * 3) # K Gaussians: 3 parameters for each (mean, variance, mixing weight) # Define a custom loss function for the negative log-likelihood mdn_loss <- function(y_true, y_pred) { k <- backend() # Get components of the MoG from the NN mus <- y_pred[, 1:(K * 1)] logsigmas <- y_pred[, (K * 1 + 1):(K * 2)] alphas_raw <- y_pred[, (K * 2 + 1):(K * 3)] alphas <- k$softmax(alphas_raw, axis = 1) # Split and reshape to match dimensions y_true_r <- k_reshape(k_repeat(y_true, K), c(nrow(y_pred)*K, 1)) mus_r <- k_reshape(mus, dim(y_true_r)) logsigmas_r <- k_reshape(logsigmas, dim(y_true_r)) sigs_r <- k$exp(logsigmas_r) alphas_r <- k_reshape(alphas, dim(y_true_r)) # calculate the negative log-likelihood of the Gaussian Mixutre terms1 = 0.5 * k$square((y_true_r - mus_r)/sigs_r) terms2 = logsigmas_r const = 0.5*k$log(2*pi) terms_nll = terms1+terms2+const loss <- k$sum(alphas_r * terms_nll) return(loss) } # Compile and train model %>% compile( loss = mdn_loss, optimizer = 'adam', metrics = list() ) training_history <- model %>% fit( x = as.matrix(X_train), y = y_train, epochs = 25, batch_size = 32, validation_split = 0.2 )
错误原因
- 计算图模式下,
y_pred的batch size是动态值(表现为None),无法用R的nrow(y_pred)获取具体数值,导致c(nrow(y_pred)*K,1)无法转换为TensorFlow可识别的张量形状 k_repeat是对整个张量进行重复,而非按元素重复,会导致维度匹配错误
修正后的代码
K <- 5 # the number of Gaussian components input_shape <- ncol(X_train) # input shape # Define the shared neural network model <- keras_model_sequential() model %>% layer_dense(units = 64, activation = "relu", input_shape = input_shape) %>% layer_dense(units = 64, activation = "relu") %>% layer_dense(units = K * 3) # K Gaussians: 3 parameters for each (mean, variance, mixing weight) # Define a custom loss function for the negative log-likelihood mdn_loss <- function(y_true, y_pred) { k <- backend() # Get components of the MoG from the NN mus <- y_pred[, 1:K] logsigmas <- y_pred[, (K+1):(2*K)] alphas_raw <- y_pred[, (2*K+1):(3*K)] alphas <- k$softmax(alphas_raw, axis = 1) # 动态构造张量形状,避免静态数值导致的类型错误 batch_size <- k$shape(y_pred)[[1]] target_shape <- k$concatenate(list(batch_size * K, 1), axis = 0) # 按元素重复y_true,匹配混合模型的分量数 y_true_r <- k$repeat_elements(y_true, rep = K, axis = 1) y_true_r <- k_reshape(y_true_r, target_shape) # 重塑其他参数张量至统一形状 mus_r <- k_reshape(mus, target_shape) logsigmas_r <- k_reshape(logsigmas, target_shape) sigs_r <- k$exp(logsigmas_r) alphas_r <- k_reshape(alphas, target_shape) # 计算负对数似然 terms1 <- 0.5 * k$square((y_true_r - mus_r) / sigs_r) terms2 <- logsigmas_r const <- 0.5 * k$log(2 * pi) terms_nll <- terms1 + terms2 + const # 对每个样本的所有分量求和,再对batch求平均(避免loss随batch size波动) loss_per_sample <- k$sum(alphas_r * terms_nll, axis = 1) loss <- k$mean(loss_per_sample) return(loss) } # Compile the model with the custom loss function model %>% compile( loss = mdn_loss, optimizer = 'adam', metrics = list() ) # Print the model summary summary(model) ### train the model training_history <- model %>% fit( x = as.matrix(X_train), # Exclude the target column y = y_train, # Target column epochs = 25, # Number of training epochs batch_size = 32, # Batch size validation_split = 0.2 # Portion of data for validation )
关键改动说明
- 动态形状构造:用
k$shape(y_pred)[[1]]获取动态batch size,通过k$concatenate构造TensorFlow可识别的形状张量,替代静态R向量 - 正确的重复操作:用
k$repeat_elements按元素重复y_true,确保每个样本标签对应K个高斯分量,替代k_repeat的整张量重复 - loss计算优化:先对每个样本的所有分量求和,再对batch求平均,避免loss值随batch size波动,训练更稳定
- 索引简化:将
y_pred的索引简化为1:K、(K+1):(2*K)等,代码可读性提升
内容的提问来源于stack exchange,提问作者fryan
相关产品推荐
相关产品推荐

