You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Azure Blob存储路径配置为DBT的Source适配Databricks?

解决dbt引用Databricks卷中CSV文件作为Source的问题

当前你的source配置会让dbt默认解析为Databricks Catalog中的物理表,但实际目标是读取卷里的CSV文件,所以需要调整配置逻辑,以下是几个可行方案:

方案一:直接在Source中定义读取文件的SQL逻辑

修改source配置,给目标表添加sql参数,直接写入read_files的查询语句,这样dbt编译时会直接使用该语句读取文件,无需依赖Catalog中的表。

示例配置:

version: 2

sources:
  - name: source_name
    tables:
      - name: table_name
        sql: "select * from read_files('<path/to/volume/folder>', format => 'csv')"

调用{{ source('source_name', 'table_name') }}时,编译后的SQL就是你原本使用的read_files查询,直接读取卷内的CSV文件。

方案二:创建Databricks外部表再映射为Source

如果需要对文件结构进行管控或优化查询性能,可以先在Databricks中创建指向卷路径的外部表,之后通过dbt source引用该表:

  1. 在Databricks执行建表语句(根据实际情况调整参数,比如是否带表头、是否自动推断 schema):
CREATE TABLE catalog_name.schema_name.table_name
USING CSV
LOCATION '/Volumes/catalog/schema/volume_name/folder_path'
OPTIONS (header = true, inferSchema = true);
  1. 修改dbt的source配置,直接引用这个已创建的外部表:
version: 2

sources:
  - name: source_name
    database: catalog_name
    schema: schema_name
    tables:
      - name: table_name

这种方式下,每日新增的符合结构的CSV文件会被外部表自动识别,适合文件格式稳定的场景。

方案三:使用新版本dbt-databricks的external_location配置

如果你使用的是dbt-databricks 1.8及以上版本,可以利用external_location参数直接配置source的文件路径,无需额外SQL:

version: 2

sources:
  - name: source_name
    tables:
      - name: table_name
        external_location: '<path/to/volume/folder>'
        file_format: csv
        options:
          header: true
          inferSchema: true

该配置会让dbt正确生成读取外部位置文件的逻辑,避免错误指向Catalog中的物理表。

内容的提问来源于stack exchange,提问作者Marc Work

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 02:13:16