You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否通过Databricks将ADLS Gen2数据导入Azure Data Explorer?

通过Databricks把ADLS Gen2数据导入ADX

前置准备

  • 服务主体得配好权限:要有ADLS Gen2的存储Blob数据读取者权限,以及ADX的数据库参与者(或能写入目标表的对应权限)
  • Databricks集群安装依赖:com.microsoft.azure.kusto:spark-kusto-connector_2.12:3.0.0(Spark版本不同的话,对应调整连接器版本,比如Spark 2.x用2.x的连接器)
  • ADX里提前建好目标表,表结构要和ADLS里的数据对应上

步骤1:用服务主体在Databricks读取ADLS数据

在Notebook里配置认证信息,读取数据:

# 替换成你的实际信息
tenant_id = "<租户ID>"
client_id = "<服务主体ID>"
client_secret = "<服务主体密钥>"
storage_account = "<ADLS存储账户名>"
container = "<容器名>"
data_path = "<数据路径,比如data/2024/*.parquet>"

# 配置Spark的ADLS认证
spark.conf.set(f"fs.azure.account.auth.type.{storage_account}.dfs.core.windows.net", "OAuth")
spark.conf.set(f"fs.azure.account.oauth.provider.type.{storage_account}.dfs.core.windows.net", "org.apache.hadoop.fs.azurebfs.oauth2.ClientCredsTokenProvider")
spark.conf.set(f"fs.azure.account.oauth2.client.id.{storage_account}.dfs.core.windows.net", client_id)
spark.conf.set(f"fs.azure.account.oauth2.client.secret.{storage_account}.dfs.core.windows.net", client_secret)
spark.conf.set(f"fs.azure.account.oauth2.client.endpoint.{storage_account}.dfs.core.windows.net", f"https://login.microsoftonline.com/{tenant_id}/oauth2/token")

# 读取数据,这里以Parquet为例,CSV/JSON的话改read.csv/read.json就行
df = spark.read.parquet(f"abfss://{container}@{storage_account}.dfs.core.windows.net/{data_path}")

步骤2:把数据写入ADX

用Kusto Spark Connector写入,同样用服务主体认证:

# 替换成你的ADX信息
adx_cluster_url = "<ADX集群地址,比如https://mycluster.eastus.kusto.windows.net>"
adx_db = "<目标数据库名>"
adx_target_table = "<目标表名>"

# 配置Kusto写入参数
kusto_config = {
    "kustoCluster": adx_cluster_url,
    "kustoDatabase": adx_db,
    "kustoTable": adx_target_table,
    "aadTenantId": tenant_id,
    "aadClientId": client_id,
    "aadClientSecret": client_secret,
    "writeMode": "Append"  # 可选Append/Create/Replace,按需选
}

# 执行写入
df.write.format("com.microsoft.kusto.spark.datasource").options(**kusto_config).save()

效率优化点

  • 按需读取:根据数据的分区字段(比如日期)过滤,只读取需要导入的部分,减少数据量
  • 调整分区:用df.repartition(10)(数字按需调整)把DataFrame分成合适的分区,每个分区大小控制在100MB-1GB,适配ADX的写入性能
  • 用列式格式:优先用Parquet/ORC这类列式存储,比CSV/JSON的读写效率高很多
  • 增量导入:如果是定期同步,记录上次导入的时间戳,下次只读取新增数据,避免重复处理

内容的提问来源于stack exchange,提问作者MMV

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 20:54:22