使用自定义损失函数的R Keras模型多数据集训练崩溃问题
问题:自定义损失函数导致Keras模型拟合第二个数据集时崩溃
我正在开发一种带有特殊损失函数的神经网络,该损失函数在项目场景中具备合理性。搭建好模型架构后,尝试将其分别拟合到数据集x_train/y_train以及x_train2/y_train2,但拟合第二个数据集时模型崩溃,报错信息如下:
Error in py_call_impl(callable, call_args$unnamed, call_args$named) : RuntimeError: in user code: File ...\DOCUME~1\VIRTUA~1\R-TENS~1\Lib\site-packages\keras\src\engine\training.py", line 1401, in train_function * return step_function(self, iterator) File ...\R\cache\R\renv\library\comoR-4eef73a7\R-4.3\x86_64-w64-mingw32\reticulate\python\rpytools\call.py", line 16, in python_function * raise error RuntimeError: NA/NaN argument
最小可复现代码
rm(list=ls()) y = rnorm(1000) x= rnorm (1000) num_classes=10 mat <- matrix( rnorm (1000*num_classes), ncol=num_classes) x_train =x y_train =mat length(x_train) dim(y_train) model <- keras_model_sequential() %>% layer_dense(units = 64, activation = 'relu', input_shape = c(1)) %>% layer_dense(units = 64, activation = 'relu' ) %>% layer_dense(units = num_classes, activation = 'softmax') custom_loss <- function(y_true, y_pred) { tt <- 0 for (i in 1:nrow(y_true)) { tt <- tt + log(sum(exp(y_true[i,]) * y_pred[i,])) } mse <- -tt return(mse) } blank_model <- model %>% compile( loss = custom_loss, optimizer = 'adam', metrics = c('accuracy') ) model1 <- blank_model model2 <- blank_model history <-model1 %>% fit( x_train, y_train, epochs = 40, batch_size = 100 ) x_train2 =x[-1] y_train2 =mat[-1,] history <-model2 %>% fit( x_train2, y_train2, epochs = 40, batch_size = 100 )
问题原因与修复方案
核心问题
- 损失函数不兼容Keras张量批次处理:原损失函数使用R原生循环和矩阵操作,但Keras训练时传递的是TensorFlow张量,
nrow(y_true)无法正确获取批次维度,循环操作会导致计算异常。 - 数值溢出产生NaN:
exp(y_true[i,])中y_true是正态分布随机值,较大的正值会让exp结果溢出为Inf,后续log计算会产生NaN。 - 模型浅拷贝干扰:
model1 <- blank_model和model2 <- blank_model是浅拷贝,两个模型共享权重,训练model1会污染model2的初始状态。
修复步骤
1. 用TensorFlow张量操作重写损失函数
替换R原生循环为TF张量操作,确保兼容批次处理,同时避免数值溢出:
custom_loss <- function(y_true, y_pred) { # 逐元素乘积 → 按样本维度求和 → 取对数 → 总和取负 product <- tf$multiply(y_true, y_pred) sum_product <- tf$reduce_sum(product, axis = 1L) log_sum <- tf$math$log(sum_product) total_loss <- -tf$reduce_sum(log_sum) return(total_loss) }
注:如果业务逻辑必须保留exp(y_true),需先对y_true做标准化处理(如tf$math$normalize(y_true)),避免数值溢出。
2. 避免模型浅拷贝,创建独立模型
通过函数封装模型创建逻辑,确保每个模型拥有独立权重:
create_model <- function() { keras_model_sequential() %>% layer_dense(units = 64, activation = 'relu', input_shape = c(1)) %>% layer_dense(units = 64, activation = 'relu') %>% layer_dense(units = num_classes, activation = 'softmax') %>% compile( loss = custom_loss, optimizer = 'adam', metrics = c('accuracy') ) } model1 <- create_model() model2 <- create_model()
3. 验证修复效果
修改后重新运行代码,拟合x_train2/y_train2时不会再出现NA/NaN错误,两个模型可独立训练不同数据集。
内容的提问来源于stack exchange,提问作者CoilyUlver
相关产品推荐
相关产品推荐

