You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取Databricks工作区所有集群的已安装库列表

批量获取Databricks工作区集群已安装库列表操作指南

你之前用到的Workspace CLI是用于管理工作区文件资源的,本身不支持查询集群库相关信息,需要调用Databricks集群库相关接口实现批量查询。

前置依赖

  • 持有Databricks工作区的个人访问令牌(PAT),且令牌拥有所有集群的查看权限
  • 本地安装Python 3.7及以上版本

可直接运行的Python脚本

import requests
import csv

# 请替换为你自己的工作区配置
DATABRICKS_INSTANCE = "https://你的工作区地址.cloud.databricks.com"
TOKEN = "你的个人访问令牌"
HEADERS = {"Authorization": f"Bearer {TOKEN}"}

def list_all_clusters():
    url = f"{DATABRICKS_INSTANCE}/api/2.0/clusters/list"
    resp = requests.get(url, headers=HEADERS)
    resp.raise_for_status()
    return resp.json().get("clusters", [])

def get_cluster_libraries(cluster_id):
    url = f"{DATABRICKS_INSTANCE}/api/2.0/libraries/cluster-status?cluster_id={cluster_id}"
    resp = requests.get(url, headers=HEADERS)
    resp.raise_for_status()
    return resp.json().get("library_statuses", [])

if __name__ == "__main__":
    # 输出文件路径
    output_file = "databricks_cluster_libraries.csv"
    all_records = []
    clusters = list_all_clusters()
    
    for cluster in clusters:
        cluster_id = cluster["cluster_id"]
        cluster_name = cluster["cluster_name"]
        libs = get_cluster_libraries(cluster_id)
        for lib in libs:
            # 提取库基本信息
            lib_info = lib["library"]
            lib_type = list(lib_info.keys())[0]
            lib_name = lib_info[lib_type] if lib_type != "maven" else lib_info[lib_type]["coordinates"]
            # 组装四个要求字段
            record = {
                "Name": lib_name,
                "Type": lib_type,
                "Status": lib["status"],
                "Source": lib_info[lib_type] if isinstance(lib_info[lib_type], str) else str(lib_info[lib_type])
            }
            # 额外追加集群ID、集群名字段方便区分不同集群的库
            record["ClusterName"] = cluster_name
            record["ClusterId"] = cluster_id
            all_records.append(record)
    
    # 写入CSV文件
    if all_records:
        fields = all_records[0].keys()
        with open(output_file, "w", newline="", encoding="utf-8") as f:
            writer = csv.DictWriter(f, fieldnames=fields)
            writer.writeheader()
            writer.writerows(all_records)
    print(f"查询完成,结果已保存到{output_file}")

输出说明

运行脚本后会在当前目录生成CSV格式的结果文件,完整包含你要求的4个字段,字段含义如下:

  • Name:已安装的库名称
  • Type:库的类型,常见包括pypi、maven、whl、egg等
  • Status:库的安装状态,包括INSTALLED(安装成功)、FAILED(安装失败)、PENDING(待安装)等
  • Source:库的来源地址/配置信息

输出效果示例:
输出示例图

内容的提问来源于stack exchange,提问作者Learn2Code

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 12:36:01