You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Azure ML Python SDK v2时Mlflow Run ID找不到,无法访问运行日志

问题

使用Azure ML jobs通过Python SDK v2运行实验,运行完成后无法访问运行日志,系统提示“run 'xxxx' not found”。确认该Run ID确实存在,且本人是工作区和集群的所有者。

代码示例

from mlflow.tracking import MlflowClient

# 使用MLflow获取刚完成的任务
run_id = 'musing_steelpan_xxxx'

finished_mlflow_run = MlflowClient().get_run(run_id)

报错信息

MlflowException                           Traceback (most recent call last)
Cell In [5], line 6
      3 # 使用MLflow获取刚完成的任务
      4 run_id = 'musing_steelpan_hnlbhxf9qy'
----> 6 finished_mlflow_run = MlflowClient().get_run(run_id)

File /miniconda/envs/benchmark/lib/python3.8/site-packages/mlflow/tracking/client.py:150, in MlflowClient.get_run(self, run_id)
    112 def get_run(self, run_id: str) -> Run:
    113     """
    114     Fetch the run from backend store. The resulting :py:class:`Run <mlflow.entities.Run>`
    115     contains a collection of run metadata -- :py:class:`RunInfo <mlflow.entities.RunInfo>`,
   (...)
    148         status: FINISHED
    149     """
--> 150     return self._tracking_client.get_run(run_id)

File /miniconda/envs/benchmark/lib/python3.8/site-packages/mlflow/tracking/_tracking_service/client.py:72, in TrackingServiceClient.get_run(self, run_id)
     58 """
     59 Fetch the run from backend store. The resulting :py:class:`Run <mlflow.entities.Run>`
     60 contains a collection of run metadata -- :py:class:`RunInfo <mlflow.entities.RunInfo>`,
  (...)
     69          raises an exception.
     70 """
     71 _validate_run_id(run_id)
   ...
    648     )
    649 run_info = self._get_run_info_from_dir(run_dir)
    650 if run_info.experiment_id != exp_id:

MlflowException: Run 'musing_steelpan_xxxx' not found
解决方案

1. 关联MLflow客户端到Azure ML工作区

默认初始化的MlflowClient可能指向本地或非目标MLflow服务器,而非你的Azure ML工作区。需显式设置Azure ML工作区的MLflow跟踪URI:

from azure.ai.ml import MLClient
from azure.identity import DefaultAzureCredential
import mlflow
from mlflow.tracking import MlflowClient

# 初始化Azure ML工作区客户端
ml_client = MLClient(
    DefaultAzureCredential(),
    subscription_id="你的订阅ID",
    resource_group_name="你的资源组名称",
    workspace_name="你的工作区名称"
)

# 获取并设置工作区的MLflow跟踪URI
tracking_uri = ml_client.workspaces.get(ml_client.workspace_name).mlflow_tracking_uri
mlflow.set_tracking_uri(tracking_uri)

# 现在获取目标Run
run_id = 'musing_steelpan_xxxx'
finished_mlflow_run = MlflowClient().get_run(run_id)

2. 确认实验上下文匹配

Azure ML中的Run隶属于特定实验,如果MLflow客户端默认的实验与目标Run所属实验不一致,也会出现找不到的情况。可以显式指定实验ID或名称:

# 方法1:设置默认实验
mlflow.set_experiment(experiment_name="你的实验名称")

# 方法2:验证Run所属实验ID
client = MlflowClient()
run = client.get_run(run_id)
print(f"Run所属实验ID: {run.info.experiment_id}")

3. 清理MLflow本地缓存

本地MLflow配置缓存可能指向旧的跟踪服务器,可按以下操作重置:

  • 删除本地mlruns目录(如果存在)
  • 重启Python环境后重新执行代码

4. 验证Run的状态与归属

在Azure ML门户中确认:

  • 目标Run确实存在于当前工作区
  • Run状态为已完成
  • Run的ID与代码中使用的完全一致(注意大小写和特殊字符)

内容的提问来源于stack exchange,提问作者Simón Cerda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 22:30:41