如何将R模型转换为原始二进制或字符串缓冲区并解决序列化异常
R模型对象序列化存储问题说明与解决方案
问题背景
在IBM Watson Studio第三方环境使用R 3.6版本时,需要将模型对象转换为原始二进制或字符串缓冲区,通过API请求传输数据完成模型文件存储。
此前方案的问题
此前参考相关技术内容,选择通过jsonlite::serializeJSON()将模型转为JSON格式再转换为raw类型,过程看似正常,但实际调用预测功能时遇到两类问题:
- 首次运行出现
could not find function "list"报错,原因是serializeJSON()无法正确转换模型对象内的复杂列表 - 针对GLM模型实施社区建议的修复方案后,又出现
object 'C_logit_linkinv' not found错误,可见该方案对不同类型的模型通用性较差
需求
当前已有的解决方案均不适用于场景,需要找到更稳定通用的、将模型对象转换为原始字节或字符串缓冲区的方法。
问题复现代码
require(jsonlite) # 生成测试数据 data <- data.frame(x1=rnorm(100), y=rbinom(100,1,.6)) # 拟合模型并执行JSON序列化转raw fitted_model <- glm(y ~ 0 + ., data = data, family = binomial) model_as_json <- jsonlite::serializeJSON(fitted_model) model_as_raw <- charToRaw(model_as_json) # 反序列化恢复模型 back_to_json <- rawToChar(model_as_raw) back_to_model <- jsonlite::unserializeJSON(back_to_json) # 测试预测 scoring_data <- data.frame(x1=rnorm(5)) predict(object=back_to_model, newdata = scoring_data, type='response')
运行环境配置
- R version 3.6.1 (2019-07-05)
- 平台:x86_64-conda_cos6-linux-gnu (64位)
- 运行系统:Red Hat Enterprise Linux 8.2 (Ootpa)
- jsonlite版本:jsonlite_1.7.1
解决方案
方案1:R原生序列化(优先选择,通用性最高)
R内置的serialize()/unserialize()函数为R对象原生序列化工具,会完整保留对象的所有属性、引用关系,支持几乎所有R模型对象,兼容性远高于JSON序列化方案。
示例代码:
# 模型直接序列化为raw二进制缓冲区 model_as_raw <- serialize(fitted_model, connection = NULL) # 反序列化恢复模型 back_to_model <- unserialize(model_as_raw) # 测试预测功能可正常运行 scoring_data <- data.frame(x1=rnorm(5)) predict(object=back_to_model, newdata = scoring_data, type='response')
如果需要转为字符串格式传输,可将raw对象编码为base64字符串,示例:
# 安装加载base64编码工具,R 3.6可兼容base64enc包 # install.packages("base64enc") require(base64enc) # raw转base64字符串 model_as_str <- base64encode(model_as_raw) # base64字符串转回raw model_as_raw2 <- base64decode(model_as_str) # 恢复模型 back_to_model2 <- unserialize(model_as_raw2)
该方案优势:
- 原生实现,无额外依赖风险,兼容性覆盖所有标准R模型对象
- 序列化速度、压缩率远高于JSON序列化方案
- 完整保留模型内部的函数引用、C语言接口指针等特殊属性,不会出现反序列化后预测失效的问题
方案2:JSON格式适配方案(仅作为API要求JSON传输时的备选)
如果API要求必须使用JSON格式传递,可在GLM模型拟合时减少冗余引用,反序列化后手动补充缺失属性,示例:
# 拟合时关闭冗余数据存储 fitted_model <- glm(y ~ 0 + ., data = data, family = binomial, model = FALSE, y = FALSE) # 序列化转JSON model_as_json <- jsonlite::serializeJSON(fitted_model) model_as_raw <- charToRaw(model_as_json) # 反序列化 back_to_json <- rawToChar(model_as_raw) back_to_model <- jsonlite::unserializeJSON(back_to_json) # 手动补充缺失的linkinv引用 back_to_model$family$linkinv <- binomial()$linkinv # 预测可正常运行 predict(object=back_to_model, newdata = scoring_data, type='response')
该方案仅针对GLM模型生效,其他类型模型需要单独适配对应缺失属性,通用性较差。
内容的提问来源于stack exchange,提问作者ricniet
相关产品推荐
相关产品推荐

