You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Databricks Asset Bundle中传递参数文件至通用Python任务

在Databricks Asset Bundle中实现通用Ingestion任务的多配置文件方案

此前在DBX中,我会创建一个通用ingestion任务,在deployment.yml中通过如下配置使用不同的YAML配置文件:

tasks:
  - task_key: ingest_<source>_<table>
    job_cluster_key: job_cluster
    python_wheel_task:
      package_name: dbx_<source>_<layer>
      entry_point: <source>_ingestion_task
      parameters:
        - --conf-file
        - file:fuse://conf/deployments/<source>/tasks/<table>_config.yml

在Python任务中,我可以通过value = self.conf["<key>"]["<sub_key>"]访问这些配置参数。
现在我需要在Databricks Asset Bundle中实现相同功能:通过独立参数文件为多个通用ingestion函数创建任务,请问该如何操作?

实现步骤

1. 组织配置文件结构

在Bundle根目录下创建和DBX一致的配置文件目录结构,比如conf/ingestions/<source>/<table>_config.yml,每个表的独立配置文件单独存放。

2. 定义通用任务模板(基于YAML锚点)

利用DAB的YAML锚点与引用特性,创建通用任务模板,避免重复编写冗余配置:

# 通用Ingestion任务锚点
common_ingest_task: &common_ingest_task
  job_cluster_key: job_cluster
  python_wheel_task:
    package_name: dbx_${source}_${layer}
    entry_point: ${source}_ingestion_task
    parameters:
      - --conf-file
      - file:fuse://conf/ingestions/${source}/${table}_config.yml

# 复用的集群配置
job_cluster: &job_cluster
  new_cluster:
    spark_version: "13.3.x-scala2.12"
    node_type_id: "Standard_DS3_v2"
    num_workers: 2

# 具体任务实例
resources:
  jobs:
    # MySQL同步user表任务
    ingest_mysql_user:
      name: "Ingest MySQL User Table"
      tasks:
        - task_key: ingest_mysql_user
          <<: *common_ingest_task
          variables:
            source: mysql
            layer: bronze
            table: user
    # PostgreSQL同步order表任务
    ingest_pg_order:
      name: "Ingest PostgreSQL Order Table"
      tasks:
        - task_key: ingest_pg_order
          <<: *common_ingest_task
          variables:
            source: postgres
            layer: bronze
            table: order

3. 批量生成任务(基于DAB模板循环)

如果需要创建大量Ingestion任务,可利用DAB的模板循环功能(要求DAB版本≥0.20.0),先定义任务清单再批量生成:

# 所有需要同步的源与表清单
ingestion_tasks:
  - source: mysql
    layer: bronze
    table: user
  - source: postgres
    layer: bronze
    table: order
  - source: s3
    layer: silver
    table: customer

resources:
  jobs:
    # 循环生成所有Ingestion任务
    {% for task in ingestion_tasks %}
    ingest_{{ task.source }}_{{ task.table }}:
      name: "Ingest {{ task.source | capitalize }} {{ task.table | capitalize }} Table"
      tasks:
        - task_key: ingest_{{ task.source }}_{{ task.table }}
          job_cluster_key: job_cluster
          python_wheel_task:
            package_name: dbx_{{ task.source }}_{{ task.layer }}
            entry_point: {{ task.source }}_ingestion_task
            parameters:
              - --conf-file
              - file:fuse://conf/ingestions/{{ task.source }}/{{ task.table }}_config.yml
    {% endfor %}

# 集群配置
job_cluster:
  new_cluster:
    spark_version: "13.3.x-scala2.12"
    node_type_id: "Standard_DS3_v2"
    num_workers: 2

4. Python任务读取配置逻辑保持不变

无需修改原有Python任务代码,依然可以通过self.conf["<key>"]["<sub_key>"]读取独立配置文件中的参数,DAB会自动传递--conf-file参数到任务中。

5. 部署Bundle

执行以下命令完成部署:

databricks bundle deploy

内容的提问来源于stack exchange,提问作者Anouar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 07:47:07