You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用自定义损失函数的R Keras模型多数据集训练崩溃问题

问题:自定义损失函数导致Keras模型拟合第二个数据集时崩溃

我正在开发一种带有特殊损失函数的神经网络,该损失函数在项目场景中具备合理性。搭建好模型架构后,尝试将其分别拟合到数据集x_train/y_train以及x_train2/y_train2,但拟合第二个数据集时模型崩溃,报错信息如下:

Error in py_call_impl(callable, call_args$unnamed, call_args$named) : 
  RuntimeError: in user code:

    File  ...\DOCUME~1\VIRTUA~1\R-TENS~1\Lib\site-packages\keras\src\engine\training.py", line 1401, in train_function  *
        return step_function(self, iterator)
    File  ...\R\cache\R\renv\library\comoR-4eef73a7\R-4.3\x86_64-w64-mingw32\reticulate\python\rpytools\call.py", line 16, in python_function  *
        raise error

    RuntimeError: NA/NaN argument

最小可复现代码

rm(list=ls())
y = rnorm(1000)
x= rnorm (1000)
num_classes=10
mat <- matrix( rnorm (1000*num_classes), ncol=num_classes)
x_train =x
y_train =mat
length(x_train)
dim(y_train)

model <- keras_model_sequential() %>%
  layer_dense(units = 64, activation = 'relu', input_shape = c(1)) %>%
  layer_dense(units = 64, activation = 'relu' ) %>%
  layer_dense(units = num_classes, activation = 'softmax')


custom_loss <- function(y_true, y_pred) {
  tt <- 0
  for (i in 1:nrow(y_true)) {
    tt <- tt + log(sum(exp(y_true[i,]) * y_pred[i,]))
  }
  mse <- -tt
  return(mse)
}
blank_model <-  model %>% compile(
  loss = custom_loss,
  optimizer = 'adam',
  metrics = c('accuracy')
)


model1 <- blank_model
model2 <- blank_model


history <-model1 %>% fit(
  x_train, y_train,
  epochs = 40,
  batch_size = 100
)

x_train2 =x[-1]
y_train2 =mat[-1,]

history <-model2 %>% fit(
  x_train2, y_train2,
  epochs = 40,
  batch_size = 100
)

问题原因与修复方案

核心问题

  1. 损失函数不兼容Keras张量批次处理:原损失函数使用R原生循环和矩阵操作,但Keras训练时传递的是TensorFlow张量,nrow(y_true)无法正确获取批次维度,循环操作会导致计算异常。
  2. 数值溢出产生NaN:exp(y_true[i,])中y_true是正态分布随机值,较大的正值会让exp结果溢出为Inf,后续log计算会产生NaN。
  3. 模型浅拷贝干扰:model1 <- blank_model和model2 <- blank_model是浅拷贝,两个模型共享权重,训练model1会污染model2的初始状态。

修复步骤

1. 用TensorFlow张量操作重写损失函数

替换R原生循环为TF张量操作,确保兼容批次处理,同时避免数值溢出:

custom_loss <- function(y_true, y_pred) {
  # 逐元素乘积 → 按样本维度求和 → 取对数 → 总和取负
  product <- tf$multiply(y_true, y_pred)
  sum_product <- tf$reduce_sum(product, axis = 1L)
  log_sum <- tf$math$log(sum_product)
  total_loss <- -tf$reduce_sum(log_sum)
  return(total_loss)
}

注:如果业务逻辑必须保留exp(y_true),需先对y_true做标准化处理(如tf$math$normalize(y_true)),避免数值溢出。

2. 避免模型浅拷贝,创建独立模型

通过函数封装模型创建逻辑,确保每个模型拥有独立权重:

create_model <- function() {
  keras_model_sequential() %>%
    layer_dense(units = 64, activation = 'relu', input_shape = c(1)) %>%
    layer_dense(units = 64, activation = 'relu') %>%
    layer_dense(units = num_classes, activation = 'softmax') %>%
    compile(
      loss = custom_loss,
      optimizer = 'adam',
      metrics = c('accuracy')
    )
}

model1 <- create_model()
model2 <- create_model()

3. 验证修复效果

修改后重新运行代码,拟合x_train2/y_train2时不会再出现NA/NaN错误,两个模型可独立训练不同数据集。


内容的提问来源于stack exchange,提问作者CoilyUlver

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 11:07:02