如何在Databricks Asset Bundle中传递参数文件至通用Python任务
在Databricks Asset Bundle中实现通用Ingestion任务的多配置文件方案
此前在DBX中,我会创建一个通用ingestion任务,在deployment.yml中通过如下配置使用不同的YAML配置文件:
tasks: - task_key: ingest_<source>_<table> job_cluster_key: job_cluster python_wheel_task: package_name: dbx_<source>_<layer> entry_point: <source>_ingestion_task parameters: - --conf-file - file:fuse://conf/deployments/<source>/tasks/<table>_config.yml在Python任务中,我可以通过
value = self.conf["<key>"]["<sub_key>"]访问这些配置参数。
现在我需要在Databricks Asset Bundle中实现相同功能:通过独立参数文件为多个通用ingestion函数创建任务,请问该如何操作?
实现步骤
1. 组织配置文件结构
在Bundle根目录下创建和DBX一致的配置文件目录结构,比如conf/ingestions/<source>/<table>_config.yml,每个表的独立配置文件单独存放。
2. 定义通用任务模板(基于YAML锚点)
利用DAB的YAML锚点与引用特性,创建通用任务模板,避免重复编写冗余配置:
# 通用Ingestion任务锚点 common_ingest_task: &common_ingest_task job_cluster_key: job_cluster python_wheel_task: package_name: dbx_${source}_${layer} entry_point: ${source}_ingestion_task parameters: - --conf-file - file:fuse://conf/ingestions/${source}/${table}_config.yml # 复用的集群配置 job_cluster: &job_cluster new_cluster: spark_version: "13.3.x-scala2.12" node_type_id: "Standard_DS3_v2" num_workers: 2 # 具体任务实例 resources: jobs: # MySQL同步user表任务 ingest_mysql_user: name: "Ingest MySQL User Table" tasks: - task_key: ingest_mysql_user <<: *common_ingest_task variables: source: mysql layer: bronze table: user # PostgreSQL同步order表任务 ingest_pg_order: name: "Ingest PostgreSQL Order Table" tasks: - task_key: ingest_pg_order <<: *common_ingest_task variables: source: postgres layer: bronze table: order
3. 批量生成任务(基于DAB模板循环)
如果需要创建大量Ingestion任务,可利用DAB的模板循环功能(要求DAB版本≥0.20.0),先定义任务清单再批量生成:
# 所有需要同步的源与表清单 ingestion_tasks: - source: mysql layer: bronze table: user - source: postgres layer: bronze table: order - source: s3 layer: silver table: customer resources: jobs: # 循环生成所有Ingestion任务 {% for task in ingestion_tasks %} ingest_{{ task.source }}_{{ task.table }}: name: "Ingest {{ task.source | capitalize }} {{ task.table | capitalize }} Table" tasks: - task_key: ingest_{{ task.source }}_{{ task.table }} job_cluster_key: job_cluster python_wheel_task: package_name: dbx_{{ task.source }}_{{ task.layer }} entry_point: {{ task.source }}_ingestion_task parameters: - --conf-file - file:fuse://conf/ingestions/{{ task.source }}/{{ task.table }}_config.yml {% endfor %} # 集群配置 job_cluster: new_cluster: spark_version: "13.3.x-scala2.12" node_type_id: "Standard_DS3_v2" num_workers: 2
4. Python任务读取配置逻辑保持不变
无需修改原有Python任务代码,依然可以通过self.conf["<key>"]["<sub_key>"]读取独立配置文件中的参数,DAB会自动传递--conf-file参数到任务中。
5. 部署Bundle
执行以下命令完成部署:
databricks bundle deploy
内容的提问来源于stack exchange,提问作者Anouar
相关产品推荐
相关产品推荐

