如何通过Databricks Asset Bundles复用集群
用Databricks Asset Bundles实现集群复用至多个作业
1. 定义独立的共享集群资源
在你的bundle.yml文件中,单独定义一个可复用的集群资源,而非嵌套在作业内部:
resources: clusters: shared_cluster: cluster_name: "shared-production-cluster" spark_version: "13.3.x-scala2.12" node_type_id: "Standard_DS3_v2" num_workers: 2 autoscale: min_workers: 2 max_workers: 5 spark_conf: spark.databricks.delta.preview.enabled: "true" # 可按需添加实例池、权限等其他配置
2. 在多个作业中引用该共享集群
定义作业时,使用existing_cluster_id字段,通过DAB的变量引用语法指向之前定义的集群资源ID:
resources: jobs: sales_etl_job: name: "Sales-ETL-Job" tasks: - task_key: "load_sales_data" existing_cluster_id: "{{ resources.clusters.shared_cluster.id }}" notebook_task: notebook_path: "/Workspace/ETL/sales/load_sales_data" inventory_report_job: name: "Inventory-Report-Job" tasks: - task_key: "generate_inventory_report" existing_cluster_id: "{{ resources.clusters.shared_cluster.id }}" spark_python_task: python_file: "/Workspace/Reports/inventory_report.py"
关键注意事项
- 共享集群的配置需兼容所有关联作业的运行要求,比如Spark版本、依赖库、资源规格等
- 若集群处于停止状态,作业执行前会自动启动集群,需考虑启动延迟对作业的影响
- 控制共享集群上的并发任务数,避免因资源耗尽导致任务失败,可通过Spark配置调整并发度
内容的提问来源于stack exchange,提问作者Sumit Desai
相关产品推荐
相关产品推荐

