You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何挂载Azure存储Gen2容器供Databricks所有Notebook访问及相关疑问

Azure存储挂载操作步骤

在Notebook中运行以下代码完成认证并创建挂载点:

configs = {"fs.azure.account.auth.type": "OAuth",
           "fs.azure.account.oauth.provider.type":       
       "org.apache.hadoop.fs.azurebfs.oauth2.ClientCredsTokenProvider",
           "fs.azure.account.oauth2.client.id": "<application-id>",
           "fs.azure.account.oauth2.client.secret": 
dbutils.secrets.get(scope="<scope-name>",key="<service-credential-key-name>"),
           "fs.azure.account.oauth2.client.endpoint": 
           "https://login.microsoftonline.com/<directory-id>/oauth2/token"}

若需要指定存储容器内的具体路径,可在源URI中添加路径后执行挂载:

dbutils.fs.mount(
    source = "abfss://<container-name>@<storage-account-name>.dfs.core.windows.net/",
    mount_point = "/mnt/<mount-name>",
    extra_configs = configs
)
技术问询解答
  • 是否需要在所有Notebook中编写上述代码,还是仅在单个Notebook中编写后调用即可?
    不需要在所有Notebook重复编写。挂载操作在集群内执行一次后,整个集群的所有Notebook都能访问该挂载点。为避免重复挂载报错,可在Notebook开头加入判断逻辑,检查挂载点是否已存在:
mount_exists = any(mount.mountPoint == "/mnt/<mount-name>" for mount in dbutils.fs.mounts())
if not mount_exists:
    # 执行挂载代码
  • 是否可将挂载代码添加至Databricks集群,实现集群启动时自动挂载?
    可以实现,推荐两种方式:

    1. 集群初始化脚本:将挂载代码编写成Shell或Python脚本,上传至DBFS或云存储,在集群配置中指定该脚本为初始化脚本,集群启动时会自动执行脚本完成挂载。
    2. 集群Spark配置:若无需显式挂载点,可直接在集群的Spark配置中添加存储账户的认证参数,实现集群级别的存储访问;但如果需要固定挂载点,初始化脚本是更合适的方案。
  • UAT与生产环境的storage-account-name不同,如何分别传入对应值?
    可通过以下几种方式实现环境区分:

    1. Databricks Secrets:为UAT和生产环境分别创建不同的密钥(如storage-uat-account、storage-prod-account),代码中根据当前环境标识读取对应密钥。
    2. 集群环境变量:在UAT和生产集群的配置中分别设置不同的环境变量(如STORAGE_ACCOUNT_NAME),代码中通过os.getenv("STORAGE_ACCOUNT_NAME")获取对应值。
    3. 参数化Notebook:将存储账户名作为Notebook参数,运行Notebook时根据环境传入对应值,适合需要手动触发挂载的场景。

内容的提问来源于stack exchange,提问作者Surender Raja

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 20:22:42