You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言Keras混合密度网络自定义损失函数报错求助

解决R Keras混合密度网络自定义损失函数的TypeError错误

问题重现

训练混合密度网络时触发如下错误:

Epoch 1/25
2023-10-23 11:55:50.657605: I tensorflow/core/grappler/optimizers/custom_graph_optimizer_registry.cc:114] Plugin optimizer for device_type GPU is enabled.
750/750 [==============================] - 3s 4ms/step - loss: 75.5268
 Error in py_call_impl(callable, call_args$unnamed, call_args$named) : 
TypeError: in user code:

File "/Users/ryanbmac/.virtualenvs/r-tensorflow/lib/python3.9/site-packages/keras/src/engine/training.py", line 1972, in test_function *
return step_function(self, iterator)
File "/Library/Frameworks/R.framework/Versions/4.2-arm64/Resources/library/reticulate/python/rpytools/call.py", line 16, in python_function *
raise error
File "/Users/ryanbmac/.virtualenvs/r-tensorflow/lib/python3.9/site-packages/keras/src/backend.py", line 3613, in reshape
return tf.reshape(x, shape)

TypeError: Failed to convert elements of (None, 1) to Tensor. Consider casting elements to a supported type. See https://www.tensorflow.org/api_docs/python/tf/dtypes for supported TF dtypes.

用户核心代码如下:

K <- 5 # the number of Gaussian components
input_shape <- ncol(X_train)  # input shape

# Define the shared neural network
model <- keras_model_sequential()
model %>%
  layer_dense(units = 64, activation = "relu", input_shape = input_shape) %>%
  layer_dense(units = 64, activation = "relu") %>%
  layer_dense(units = K * 3)  # K Gaussians: 3 parameters for each (mean, variance, mixing weight)

# Define a custom loss function for the negative log-likelihood
mdn_loss <- function(y_true, y_pred) {
  k <- backend()
  
  # Get components of the MoG from the NN
  mus <- y_pred[, 1:(K * 1)]
  logsigmas <- y_pred[, (K * 1 + 1):(K * 2)]
  alphas_raw <- y_pred[, (K * 2 + 1):(K * 3)]
  alphas <- k$softmax(alphas_raw, axis = 1)
  
  # Split and reshape to match dimensions
  y_true_r <- k_reshape(k_repeat(y_true, K), c(nrow(y_pred)*K, 1))
  mus_r <- k_reshape(mus, dim(y_true_r))
  logsigmas_r <- k_reshape(logsigmas, dim(y_true_r))
  sigs_r <- k$exp(logsigmas_r)
  alphas_r <- k_reshape(alphas, dim(y_true_r))
    
  # calculate the negative log-likelihood of the Gaussian Mixutre
  terms1 = 0.5 * k$square((y_true_r - mus_r)/sigs_r) 
  terms2 = logsigmas_r
  const = 0.5*k$log(2*pi)
  terms_nll = terms1+terms2+const
  loss <- k$sum(alphas_r * terms_nll)
  return(loss)
}

# Compile and train
model %>%
  compile(
    loss = mdn_loss,
    optimizer = 'adam',           
    metrics = list()
  )

training_history <- 
  model %>%
  fit(
    x = as.matrix(X_train),
    y = y_train,
    epochs = 25,
    batch_size = 32,
    validation_split = 0.2
  )

错误原因

  1. 计算图模式下,y_pred的batch size是动态值(表现为None),无法用R的nrow(y_pred)获取具体数值,导致c(nrow(y_pred)*K,1)无法转换为TensorFlow可识别的张量形状
  2. k_repeat是对整个张量进行重复,而非按元素重复,会导致维度匹配错误

修正后的代码

K <- 5 # the number of Gaussian components
input_shape <- ncol(X_train)  # input shape

# Define the shared neural network
model <- keras_model_sequential()
model %>%
  layer_dense(units = 64, activation = "relu", input_shape = input_shape) %>%
  layer_dense(units = 64, activation = "relu") %>%
  layer_dense(units = K * 3)  # K Gaussians: 3 parameters for each (mean, variance, mixing weight)

# Define a custom loss function for the negative log-likelihood
mdn_loss <- function(y_true, y_pred) {
  k <- backend()
  
  # Get components of the MoG from the NN
  mus <- y_pred[, 1:K]
  logsigmas <- y_pred[, (K+1):(2*K)]
  alphas_raw <- y_pred[, (2*K+1):(3*K)]
  alphas <- k$softmax(alphas_raw, axis = 1)
  
  # 动态构造张量形状,避免静态数值导致的类型错误
  batch_size <- k$shape(y_pred)[[1]]
  target_shape <- k$concatenate(list(batch_size * K, 1), axis = 0)
  
  # 按元素重复y_true,匹配混合模型的分量数
  y_true_r <- k$repeat_elements(y_true, rep = K, axis = 1)
  y_true_r <- k_reshape(y_true_r, target_shape)
  
  # 重塑其他参数张量至统一形状
  mus_r <- k_reshape(mus, target_shape)
  logsigmas_r <- k_reshape(logsigmas, target_shape)
  sigs_r <- k$exp(logsigmas_r)
  alphas_r <- k_reshape(alphas, target_shape)
    
  # 计算负对数似然
  terms1 <- 0.5 * k$square((y_true_r - mus_r) / sigs_r) 
  terms2 <- logsigmas_r
  const <- 0.5 * k$log(2 * pi)
  terms_nll <- terms1 + terms2 + const
  # 对每个样本的所有分量求和,再对batch求平均(避免loss随batch size波动)
  loss_per_sample <- k$sum(alphas_r * terms_nll, axis = 1)
  loss <- k$mean(loss_per_sample)
  return(loss)
}

# Compile the model with the custom loss function
model %>%
  compile(
    loss = mdn_loss,
    optimizer = 'adam',           
    metrics = list()
  )

# Print the model summary
summary(model)

### train the model
training_history <- 
  model %>%
  fit(
    x = as.matrix(X_train),  # Exclude the target column
    y = y_train,             # Target column
    epochs = 25,                   # Number of training epochs
    batch_size = 32,               # Batch size
    validation_split = 0.2         # Portion of data for validation
  )

关键改动说明

  1. 动态形状构造:用k$shape(y_pred)[[1]]获取动态batch size,通过k$concatenate构造TensorFlow可识别的形状张量,替代静态R向量
  2. 正确的重复操作:用k$repeat_elements按元素重复y_true,确保每个样本标签对应K个高斯分量,替代k_repeat的整张量重复
  3. loss计算优化:先对每个样本的所有分量求和,再对batch求平均,避免loss值随batch size波动,训练更稳定
  4. 索引简化:将y_pred的索引简化为1:K、(K+1):(2*K)等,代码可读性提升

内容的提问来源于stack exchange,提问作者fryan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 18:47:15