You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何备份Azure Databricks工作区并实现Python/Terraform自动化?

Azure Databricks工作区完整备份自动化方案(Python/Terraform)

要实现Azure Databricks工作区的灾备级完整备份,需覆盖笔记本、集群配置、作业、MLflow实验、权限、库等核心组件,以下是Python和Terraform的具体实现方案:

一、Python自动化备份(基于Databricks SDK)

官方推荐使用databricks-sdk替代旧版REST API封装,先安装依赖:

pip install databricks-sdk mlflow

1. 笔记本备份(递归导出所有路径)

from databricks.sdk import WorkspaceClient
from databricks.sdk.service import workspace
import os

# 初始化Workspace客户端(自动读取环境变量或~/.databrickscfg的认证信息)
w = WorkspaceClient()

def export_notebooks(source_path: str, dest_root: str):
    """递归导出Databricks工作区中的笔记本到本地目录"""
    for item in w.workspace.list(source_path):
        local_path = f"{dest_root}{item.path}"
        if item.is_dir:
            os.makedirs(local_path, exist_ok=True)
            export_notebooks(item.path, dest_root)
        elif item.path.endswith(('.py', '.ipynb', '.scala', '.r')):
            # 以源码格式导出笔记本
            content = w.workspace.export(item.path, format=workspace.ExportFormat.SOURCE)
            with open(local_path, 'wb') as f:
                f.write(content.content)

# 导出所有用户笔记本到本地备份目录
export_notebooks("/Users", "./databricks_backups/notebooks")

2. 作业配置备份

import json

def export_jobs(dest_file: str):
    """导出所有作业的完整配置到JSON文件"""
    jobs = list(w.jobs.list())
    job_configs = [job.as_dict() for job in jobs]
    with open(dest_file, 'w') as f:
        json.dump(job_configs, f, indent=2)

export_jobs("./databricks_backups/jobs.json")

3. 集群配置备份

def export_clusters(dest_file: str):
    """导出所有集群(含已终止)的配置到JSON文件"""
    clusters = list(w.clusters.list())
    cluster_configs = [cluster.as_dict() for cluster in clusters]
    with open(dest_file, 'w') as f:
        json.dump(cluster_configs, f, indent=2)

export_clusters("./databricks_backups/clusters.json")

4. MLflow实验与运行备份

from mlflow.tracking import MlflowClient

def export_mlflow_experiments(dest_root: str):
    """导出MLflow实验元数据及运行数据"""
    mlflow_client = MlflowClient()
    experiments = mlflow_client.list_experiments()
    
    for exp in experiments:
        exp_dir = f"{dest_root}/mlflow/{exp.experiment_id}"
        os.makedirs(exp_dir, exist_ok=True)
        
        # 导出实验元数据
        with open(f"{exp_dir}/experiment_metadata.json", 'w') as f:
            json.dump(exp.__dict__, f, indent=2)
        
        # 导出所有运行的参数、指标、标签
        runs = mlflow_client.search_runs(exp.experiment_id)
        runs_data = [{
            "run_id": run.info.run_id,
            "params": run.data.params,
            "metrics": run.data.metrics,
            "tags": run.data.tags
        } for run in runs]
        with open(f"{exp_dir}/runs.json", 'w') as f:
            json.dump(runs_data, f, indent=2)

export_mlflow_experiments("./databricks_backups")

5. 权限备份(可选)

def export_permissions(dest_file: str):
    """导出工作区资源的权限配置"""
    resources = [
        {"type": "workspace", "path": "/"},
        {"type": "cluster", "cluster_id": "*"},
        {"type": "job", "job_id": "*"}
    ]
    permissions = []
    for res in resources:
        if res["type"] == "workspace":
            perm = w.permissions.get(res["type"], path=res["path"])
        elif res["type"] == "cluster":
            perm = w.permissions.get(res["type"], cluster_id=res["cluster_id"])
        elif res["type"] == "job":
            perm = w.permissions.get(res["type"], job_id=res["job_id"])
        permissions.append(perm.as_dict())
    
    with open(dest_file, 'w') as f:
        json.dump(permissions, f, indent=2)

export_permissions("./databricks_backups/permissions.json")

二、Terraform自动化备份与恢复

Terraform适合将Databricks工作区资源以基础设施即代码(IaC)的形式备份,同时支持一键恢复:

1. 配置Databricks Provider

创建provider.tf:

terraform {
  required_providers {
    databricks = {
      source  = "databricks/databricks"
      version = ">= 1.20.0"
    }
  }
}

provider "databricks" {
  # 认证方式:使用Azure CLI认证(需提前登录az login)
  azure_workspace_resource_id = "/subscriptions/<你的订阅ID>/resourceGroups/<资源组名>/providers/Microsoft.Databricks/workspaces/<工作区名>"
}

2. 导出现有工作区资源到Terraform文件

使用Databricks官方导出工具:

# 安装导出工具
go install github.com/databricks/terraform-provider-databricks/cmd/databricks_export@latest

# 导出所有资源到指定目录
databricks_export --output-dir ./databricks_terraform_backup

该工具会自动生成笔记本、作业、集群、库、权限等资源的.tf配置文件,可直接用于恢复工作区。

3. 恢复工作区

将导出的Terraform文件复制到目标环境,执行:

terraform init
terraform apply

关键注意事项

  • Delta表备份:Databricks仅备份Delta表的元数据,实际数据存储在云存储(ADLS Gen2等),需单独配置云提供商的备份方案(如Azure Blob快照、生命周期管理)。
  • 自动化调度:Python脚本可封装为Databricks Job或Azure Function,定期执行备份;Terraform导出可集成到GitHub Actions/Azure DevOps流水线,将配置文件推送到版本控制仓库。
  • 恢复测试:定期从备份中恢复到测试环境,验证备份完整性与恢复流程有效性。

内容的提问来源于stack exchange,提问作者Thabsheer Hussain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 12:22:47