Keras R自定义指标:计算y_pred前1%样本对应y_true的均值
问题核心原因
你之前使用的k_in_top_k是分类任务专用接口,作用是判断真实类别是否属于预测概率的前K个,返回布尔张量,完全不符合筛选y_pred前K个样本对应y_true的需求,这是训练无法启动的核心原因。
方案1:y_pred排名前1%对应y_true均值(优先实现)
直接基于批次大小动态计算Top 1%的样本量,自动适配不同batch size:
library(keras) Metric_Top_1Pct_Rets <- custom_metric("Top_1Pct_Rets", function(y_true, y_pred) { # 获取当前批次样本量 batch_size <- k_cast(k_shape(y_pred)[1], dtype = "float32") # 计算前1%样本量,最少取1个避免小batch报错 top_k <- k_cast(k_maximum(c(1, batch_size * 0.01)), dtype = "int32") # 获取y_pred排名前top_k的样本索引 top_indices <- k_top_k(k_flatten(y_pred), k = top_k)$indices # 提取对应位置的真实标签 top_y_true <- k_gather(k_flatten(y_true), top_indices) # 返回均值 return(k_mean(top_y_true)) }) # 编译模型时调用 model1 %>% compile( loss = "mse", optimizer = optimizer_adam(learning_rate = 0.0001), metrics = list(Metric_Top_1Pct_Rets) )
方案2:固定数量Top N样本对应y_true均值
适配你代码中取前1000个的需求,自动兼容batch size不足1000的场景:
Metric_Top_1000_Rets <- custom_metric("Top_1000_Rets", function(y_true, y_pred) { fixed_top_k <- 1000L # 取固定值和批次大小的最小值,避免索引越界 top_k <- k_minimum(c(fixed_top_k, k_cast(k_shape(y_pred)[1], dtype = "int32"))) top_indices <- k_top_k(k_flatten(y_pred), k = top_k)$indices top_y_true <- k_gather(k_flatten(y_true), top_indices) return(k_mean(top_y_true)) })
方案3:y_pred大于固定阈值对应y_true均值
按阈值筛选符合条件的样本,自动处理无符合样本的边界情况:
Metric_Threshold_Rets <- custom_metric("Threshold_Rets", function(y_true, y_pred) { threshold <- 0.8 # 可自定义阈值 mask <- k_greater(k_flatten(y_pred), threshold) filtered_y_true <- k_flatten(y_true)[mask] # 无符合样本时返回0,可根据需求调整默认值 return(k_mean(k_switch(k_any(mask), filtered_y_true, k_constant(0, dtype = k_floatx())))) })
注意事项
- 所有运算必须使用Keras后端
k_开头的函数,禁止使用R原生的mean()、sort()等函数,否则静态计算图构建失败会导致训练无法启动 - 若你的y_true/y_pred为二维及以上张量,可根据数据结构调整
k_flatten()的逻辑,保证筛选维度对应样本维度 - 小batch训练时前1%样本量过少会导致指标波动较大,建议batch size设置不小于1000
内容的提问来源于stack exchange,提问作者Chrispy1098
相关产品推荐
相关产品推荐

