You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python列出Azure Databricks中Azure VM托管身份创建的作业运行?遇403错误

Python解决方案:列出Azure Databricks托管身份创建的作业运行

问题根源

你的HTTP 403错误核心原因有两个:

  • 令牌资源范围错误:你请求的是Azure管理平台的令牌(management.azure.com/),但Databricks API需要的是专属的Databricks服务资源令牌
  • API端点错误:你调用的是集群列表接口(/api/2.0/clusters/list),而非作业运行列表接口

修正后的实现代码

import requests
from azure.identity import DefaultAzureCredential

# 初始化凭据(VM托管身份会被DefaultAzureCredential自动识别)
credential = DefaultAzureCredential()

# 获取Databricks API专属访问令牌,资源ID为固定值
access_token = credential.get_token("2ff814a6-3304-4ab8-85cb-cd0e6f879c1d/.default").token

# 配置Databricks工作区地址与作业运行列表API端点
databricks_workspace_url = "https://<你的Databricks工作区URL>"
api_endpoint = f"{databricks_workspace_url}/api/2.1/jobs/runs/list"

# 设置请求头
headers = {
    "Authorization": f"Bearer {access_token}",
    "Content-Type": "application/json"
}

# 可选:添加筛选条件(如指定作业ID、运行状态),需用POST请求
# payload = {
#     "job_id": <目标作业ID>,
#     "statuses": ["SUCCEEDED", "FAILED"]
# }
# response = requests.post(api_endpoint, headers=headers, json=payload)

# 发送GET请求获取所有作业运行
response = requests.get(api_endpoint, headers=headers)

# 处理响应结果
if response.status_code == 200:
    runs_data = response.json()
    print("作业运行列表:")
    for run in runs_data.get("runs", []):
        print(f"运行ID: {run['run_id']}, 作业ID: {run['job_id']}, 状态: {run['state']['life_cycle_state']}")
else:
    print(f"请求失败,状态码: {response.status_code}")
    print(f"错误详情: {response.json()}")

必备权限配置

要彻底解决403问题,需确保VM托管身份拥有以下权限:

  • 在Databricks工作区中,被分配作业查看者(Job Viewer)或更高层级的角色
  • 若使用Azure RBAC控制,需为托管身份分配Databricks工作区参与者或对应的数据访问权限

关键细节说明

  • DefaultAzureCredential会自动优先使用VM的托管身份,无需额外手动配置
  • 资源ID 2ff814a6-3304-4ab8-85cb-cd0e6f879c1d 是Azure全局范围内的Databricks服务固定标识,必须准确填写
  • 作业运行列表API支持通过POST请求传递筛选参数,可按作业ID、运行状态、时间范围等条件过滤结果

内容的提问来源于stack exchange,提问作者user3651363

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 03:02:02