Azure ML中如何生成model.pkl及转换data.ilearner为该格式?
Azure ML中Boosted Decision Tree Regression模型转model.pkl的解决方案
一、训练时直接生成model.pkl文件
场景1:使用AutoML训练
AutoML训练完成后,可通过SDK直接导出或下载pickle格式的模型:
from azureml.core import Workspace, Experiment, Run import joblib # 加载工作区(需提前配置好config.json) ws = Workspace.from_config() # 指定实验和Run ID experiment = Experiment(ws, "你的实验名称") automl_run = Run(experiment, "目标Run的ID") # 获取最佳模型并保存为pkl best_run, fitted_model = automl_run.get_output() joblib.dump(fitted_model, "model.pkl") # 可选:将pkl文件上传到Azure ML工作区 automl_run.upload_file("outputs/model.pkl", "model.pkl") # 或者直接从Run下载已生成的pkl automl_run.download_file("outputs/model.pkl", "本地保存路径/model.pkl")
场景2:自定义训练脚本使用内置算法
在训练脚本中,训练完成后直接用joblib保存模型对象,替代Azure ML自动生成的data.ilearner:
from azureml.core import Experiment, Workspace from azureml.train.automl import AutoMLConfig import pandas as pd from sklearn.model_selection import train_test_split import joblib ws = Workspace.from_config() experiment = Experiment(ws, "boosted_tree_reg_exp") # 加载并拆分数据 df = pd.read_csv("训练数据.csv") X_train, X_test, y_train, y_test = train_test_split( df.drop("目标列名", axis=1), df["目标列名"], test_size=0.2 ) # 配置AutoML,指定仅使用Boosted Decision Tree算法 automl_config = AutoMLConfig( task="regression", primary_metric="r2_score", training_data=df, label_column_name="目标列名", allowed_models=["BoostedDecisionTree"], n_cross_validations=5 ) # 提交训练 run = experiment.submit(automl_config) run.wait_for_completion() # 保存最佳模型为pkl best_run, fitted_model = run.get_output() joblib.dump(fitted_model, "model.pkl") run.upload_file("outputs/model.pkl", "model.pkl")
二、将已有的data.ilearner转换为model.pkl
data.ilearner是Azure ML的封装格式,包含模型元数据、预处理逻辑等,无法直接用LightGBM原生方法加载。需通过Azure ML SDK加载后再转存为pickle:
from azureml.core import Workspace, Model from azureml.train.automl.runtime.automl_explain_utilities import load_model import joblib # 加载工作区 ws = Workspace.from_config() # 方式1:从Azure ML工作区加载已注册的模型 model = Model(ws, name="你的模型名称") model.download(target_dir="./", exist_ok=True) # 方式2:直接加载本地的data.ilearner文件(跳过注册步骤) # 确保当前目录下有data.ilearner文件 # 加载模型 fitted_model = load_model("./data.ilearner") # 保存为pkl格式 joblib.dump(fitted_model, "model.pkl")
错误原因说明
你用lgb.Booster(model_file="data.ilearner")加载模型失败,是因为data.ilearner不是LightGBM原生的模型格式,而是Azure ML对模型的封装文件,必须使用Azure ML的专用工具加载,不能直接用LightGBM的API读取。
内容的提问来源于stack exchange,提问作者THM
相关产品推荐
相关产品推荐

