使用mlflow run运行项目时,如何让Python脚本指定的experiment-id生效?
MLflow通过MLproject运行时实验不匹配的解决方法
我用MLflow在Python脚本中跟踪运行记录,脚本内已编写实验检查与指定逻辑,确保运行记录归属目标实验:
mlflow.set_tracking_uri("http://127.0.0.1:5000") exps = mlflow.search_experiments() exp_names = [e.name for e in exps] if EXP_NAME in exp_names: experiment_id = exps[exp_names.index(EXP_NAME)].experiment_id else: experiment_id = mlflow.create_experiment(EXP_NAME) print(f"Created New Experiment > ID: {experiment_id}") experiment = mlflow.set_experiment(experiment_id=experiment_id) with mlflow.start_run(run_name="fantastic-run", experiment_id=experiment_id): # 运行逻辑 mlflow.endrun()
直接用python my_python_script.py执行脚本一切正常,但通过MLproject的my-entry条目,用mlflow run -e my-entry path_to_my_mlflow_project/执行时,系统会默认使用ID=0的实验,且脚本抛出错误:
mlflow.exceptions.MlflowException: Cannot start run with ID 4552f1852c0d4eafae72b84d38af7346 because active run ID does not match environment run ID. Make sure --experiment-name or --experiment-id matches experiment set with set_experiment(), or just use command-line arguments ERROR mlflow.cli: === Run (ID '4552f1852c0d4eafae72b84d38af7346') failed ===
解决方法
命令行直接指定实验
在mlflow run命令中添加--experiment-name或--experiment-id参数,与脚本内的EXP_NAME保持一致:mlflow run -e my-entry path_to_my_mlflow_project/ --experiment-name "你的实验名称"这样MLflow会直接使用指定实验,避免脚本内的实验设置与命令行启动的运行冲突。
修改脚本适配MLproject运行场景
用mlflow run启动时,MLflow会自动创建活跃运行,此时脚本手动调用mlflow.start_run会触发冲突。可以先检查是否存在活跃运行,再决定是否创建新运行:mlflow.set_tracking_uri("http://127.0.0.1:5000") exps = mlflow.search_experiments() exp_names = [e.name for e in exps] if EXP_NAME in exp_names: experiment_id = exps[exp_names.index(EXP_NAME)].experiment_id else: experiment_id = mlflow.create_experiment(EXP_NAME) print(f"Created New Experiment > ID: {experiment_id}") # 检查当前是否有活跃运行,无则启动新运行,有则复用 active_run = mlflow.active_run() if not active_run: with mlflow.start_run(run_name="fantastic-run", experiment_id=experiment_id): # 你的运行逻辑代码 pass else: with mlflow.start_run(run_id=active_run.info.run_id, experiment_id=experiment_id): # 你的运行逻辑代码 pass在MLproject文件中配置实验参数
可以在MLproject的entry点中定义实验参数,让脚本读取命令行传入的值:
首先修改MLproject文件:name: my-project entry_points: my-entry: command: "python my_python_script.py --experiment-name {experiment_name}" parameters: experiment_name: type: string default: "默认实验名称"然后修改脚本,添加参数解析:
import argparse parser = argparse.ArgumentParser() parser.add_argument("--experiment-name", type=str, default="默认实验名称") args = parser.parse_args() EXP_NAME = args.experiment_name # 原有的实验检查逻辑保持不变 mlflow.set_tracking_uri("http://127.0.0.1:5000") exps = mlflow.search_experiments() exp_names = [e.name for e in exps] if EXP_NAME in exp_names: experiment_id = exps[exp_names.index(EXP_NAME)].experiment_id else: experiment_id = mlflow.create_experiment(EXP_NAME) print(f"Created New Experiment > ID: {experiment_id}") # 复用活跃运行或启动新运行的逻辑同上运行时通过
-P传递实验名称:mlflow run -e my-entry path_to_my_mlflow_project/ -P experiment-name="你的实验名称"
内容的提问来源于stack exchange,提问作者inarighas
相关产品推荐
相关产品推荐

