如何保存并重新读取基于Parsnip/Agua的H2O模型对象?
问题:使用agua包保存/加载模型后的预测与模型排名问题
我用tidymodels的agua包编写了如下脚本:
library(tidymodels) library(agua) library(ggplot2) theme_set(theme_bw()) h2o_start() data(concrete) set.seed(4595) concrete_split <- initial_split(concrete, strata = compressive_strength) concrete_train <- training(concrete_split) concrete_test <- testing(concrete_split) # 最长运行120秒 auto_spec <- auto_ml() %>% set_engine("h2o", max_runtime_secs = 120, seed = 1) %>% set_mode("regression") normalized_rec <- recipe(compressive_strength ~ ., data = concrete_train) %>% step_normalize(all_predictors()) auto_wflow <- workflow() %>% add_model(auto_spec) %>% add_recipe(normalized_rec) auto_fit <- fit(auto_wflow, data = concrete_train) saveRDS(auto_fit, file = "test.h2o.auto_fit.rds") # 保存对象 h2o_end()
尝试读取保存的auto_fit对象并预测测试数据时,出现错误:
h2o_start() auto_fit <- readRDS("test.h2o.auto_fit.rds") predict(auto_fit, new_data = concrete_test)
错误信息:
Error in `h2o_get_model()`: ! Model id does not exist on the h2o server.
期望的预测结果:
predict(auto_fit, new_data = concrete_test) #> # A tibble: 260 × 1 #> .pred #> <dbl> #> 1 40.0 #> 2 43.0 #> 3 38.2 #> 4 55.7 #> 5 41.4 #> 6 28.1 #> 7 53.2 #> 8 34.5 #> 9 51.1 #> 10 37.9 #> # … with 250 more rows
更新:使用bundle方法后的新问题
按照建议改用bundle处理后:
auto_fit <- fit(auto_wflow, data = concrete_train) auto_fit_bundle <- bundle(auto_fit) saveRDS(auto_fit_bundle, file = "test.h2o.auto_fit.rds") # 保存对象 h2o_end() # 重新加载 h2o_start() auto_fit_bundle <- readRDS("test.h2o.auto_fit.rds") auto_fit <- unbundle(auto_fit_bundle) predict(auto_fit, new_data = concrete_test) rank_results(auto_fit)
此时predict可以正常运行,但调用rank_results时出错:
Error in UseMethod("rank_results") : no applicable method for 'rank_results' applied to an object of class "c('H2ORegressionModel', 'H2OModel', 'Keyed')"
解决方法
1. 直接保存模型报错的原因与解决
直接用saveRDS保存agua训练的模型时,仅保存了模型的元数据(如模型ID),实际H2O模型对象存储在H2O服务器中。重启H2O服务器后,原模型ID对应的模型已不存在,因此predict报错。
使用bundle包序列化整个模型对象(包含H2O模型本身)是正确做法,这也是你后续尝试的方法,可解决预测问题。
2. rank_results报错的解决
unbundle后,auto_fit不再是agua的auto_ml工作流拟合对象,而是纯H2O模型对象,rank_results是tidymodels针对auto_ml拟合结果设计的函数,无法直接处理纯H2O模型。可通过两种方式解决:
方法一:保存模型时同步保存排名结果
训练完成后、保存bundle前,先调用rank_results并保存结果:
auto_fit <- fit(auto_wflow, data = concrete_train) # 获取并保存排名结果 rank_res <- rank_results(auto_fit) saveRDS(rank_res, file = "test.h2o.rank_results.rds") # 保存bundle后的模型 auto_fit_bundle <- bundle(auto_fit) saveRDS(auto_fit_bundle, file = "test.h2o.auto_fit.rds") h2o_end()
后续加载时直接读取保存的排名结果:
rank_res <- readRDS("test.h2o.rank_results.rds") print(rank_res)
方法二:从H2O模型对象中获取原生排名信息
若已用bundle保存模型,unbundle后可通过H2O原生函数h2o.get_leaderboard获取AutoML模型排名:
# auto_fit为unbundle后的H2O模型 # 获取对应的AutoML对象 aml <- h2o.getModel(auto_fit@model_id$name) # 获取leaderboard leaderboard <- aml@leaderboard # 转换为tidy格式 rank_res <- as_tibble(leaderboard) print(rank_res)
这样即可得到类似rank_results的模型排名结果。
内容的提问来源于stack exchange,提问作者littleworth
相关产品推荐
相关产品推荐

