使用Azure ML Python SDK v2时Mlflow Run ID找不到,无法访问运行日志
问题
使用Azure ML jobs通过Python SDK v2运行实验,运行完成后无法访问运行日志,系统提示“run 'xxxx' not found”。确认该Run ID确实存在,且本人是工作区和集群的所有者。
代码示例
from mlflow.tracking import MlflowClient # 使用MLflow获取刚完成的任务 run_id = 'musing_steelpan_xxxx' finished_mlflow_run = MlflowClient().get_run(run_id)
报错信息
MlflowException Traceback (most recent call last) Cell In [5], line 6 3 # 使用MLflow获取刚完成的任务 4 run_id = 'musing_steelpan_hnlbhxf9qy' ----> 6 finished_mlflow_run = MlflowClient().get_run(run_id) File /miniconda/envs/benchmark/lib/python3.8/site-packages/mlflow/tracking/client.py:150, in MlflowClient.get_run(self, run_id) 112 def get_run(self, run_id: str) -> Run: 113 """ 114 Fetch the run from backend store. The resulting :py:class:`Run <mlflow.entities.Run>` 115 contains a collection of run metadata -- :py:class:`RunInfo <mlflow.entities.RunInfo>`, (...) 148 status: FINISHED 149 """ --> 150 return self._tracking_client.get_run(run_id) File /miniconda/envs/benchmark/lib/python3.8/site-packages/mlflow/tracking/_tracking_service/client.py:72, in TrackingServiceClient.get_run(self, run_id) 58 """ 59 Fetch the run from backend store. The resulting :py:class:`Run <mlflow.entities.Run>` 60 contains a collection of run metadata -- :py:class:`RunInfo <mlflow.entities.RunInfo>`, (...) 69 raises an exception. 70 """ 71 _validate_run_id(run_id) ... 648 ) 649 run_info = self._get_run_info_from_dir(run_dir) 650 if run_info.experiment_id != exp_id: MlflowException: Run 'musing_steelpan_xxxx' not found
解决方案
1. 关联MLflow客户端到Azure ML工作区
默认初始化的MlflowClient可能指向本地或非目标MLflow服务器,而非你的Azure ML工作区。需显式设置Azure ML工作区的MLflow跟踪URI:
from azure.ai.ml import MLClient from azure.identity import DefaultAzureCredential import mlflow from mlflow.tracking import MlflowClient # 初始化Azure ML工作区客户端 ml_client = MLClient( DefaultAzureCredential(), subscription_id="你的订阅ID", resource_group_name="你的资源组名称", workspace_name="你的工作区名称" ) # 获取并设置工作区的MLflow跟踪URI tracking_uri = ml_client.workspaces.get(ml_client.workspace_name).mlflow_tracking_uri mlflow.set_tracking_uri(tracking_uri) # 现在获取目标Run run_id = 'musing_steelpan_xxxx' finished_mlflow_run = MlflowClient().get_run(run_id)
2. 确认实验上下文匹配
Azure ML中的Run隶属于特定实验,如果MLflow客户端默认的实验与目标Run所属实验不一致,也会出现找不到的情况。可以显式指定实验ID或名称:
# 方法1:设置默认实验 mlflow.set_experiment(experiment_name="你的实验名称") # 方法2:验证Run所属实验ID client = MlflowClient() run = client.get_run(run_id) print(f"Run所属实验ID: {run.info.experiment_id}")
3. 清理MLflow本地缓存
本地MLflow配置缓存可能指向旧的跟踪服务器,可按以下操作重置:
- 删除本地
mlruns目录(如果存在) - 重启Python环境后重新执行代码
4. 验证Run的状态与归属
在Azure ML门户中确认:
- 目标Run确实存在于当前工作区
- Run状态为已完成
- Run的ID与代码中使用的完全一致(注意大小写和特殊字符)
内容的提问来源于stack exchange,提问作者Simón Cerda
相关产品推荐
相关产品推荐

