You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Synapse如何读取含多类型文档的Azure Cosmos DB容器指定类型数据

根因说明

Cosmos DB Spark 连接器默认开启 schema 推断时,仅会对拉取到的前N条文档做采样来生成schema,如果你的特定类型文档不在采样范围内、或是容器内混存多类型异构文档,就会出现schema仅适配部分文档的问题。

解决方案

以下三种方式按优先级从高到低排列:

方式1:自定义查询过滤特定类型文档后推断schema(推荐)

该方案可以直接在Cosmos DB侧完成数据过滤,只拉取你需要的目标类型文档,后续schema推断也只会基于目标类型的文档生成,性能和准确性都最优。
前提是你的文档有区分类型的标识字段(比如通用的type字段,或是某类文档独有的字段),示例代码如下:

cfg = {
    "spark.cosmos.accountEndpoint": Endpoint,
    "spark.cosmos.accountKey": accountKey,
    "spark.cosmos.database": databaseName,
    "spark.cosmos.container": containerName,
    # 替换为你自己的过滤规则,比如按type字段筛选、或是按独有字段存在性筛选
    "spark.cosmos.read.customQuery": "SELECT * FROM c WHERE c.type = '你需要的文档类型值'",
    "spark.cosmos.read.inferSchema.enabled": "true"
}

df = spark.read.format("cosmos.oltp").options(**cfg).load()

方式2:扩大schema推断的采样范围

如果需要读取全量文档同时保证schema覆盖所有字段,可以调整采样参数,让连接器采样更多文档来生成schema:

cfg = {
    "spark.cosmos.accountEndpoint": Endpoint,
    "spark.cosmos.accountKey": accountKey,
    "spark.cosmos.database": databaseName,
    "spark.cosmos.container": containerName,
    "spark.cosmos.read.inferSchema.enabled": "true",
    # 采样数量设为-1代表全量采样,数据量较大时会增加schema推断的耗时
    "spark.cosmos.read.inferSchema.sampleSize": "-1"
}

df = spark.read.format("cosmos.oltp").options(**cfg).load()

如果需要单独筛选特定类型,可以在读出df后再加过滤条件df = df.filter(df["type"] == "目标类型"),但该方式需要先拉取全量容器数据,性能低于方式1。

方式3:显式指定目标类型schema(稳定性最高)

如果目标类型的文档结构是固定的,直接手动定义schema可以完全避免推断误差,性能也最优:

from pyspark.sql.types import StructType, StructField, StringType, IntegerType, BooleanType

# 替换为目标类型文档对应的实际字段、字段类型
target_schema = StructType([
    StructField("id", StringType(), nullable=True),
    StructField("name", StringType(), nullable=True),
    StructField("age", IntegerType(), nullable=True),
    StructField("is_valid", BooleanType(), nullable=True)
])

cfg = {
    "spark.cosmos.accountEndpoint": Endpoint,
    "spark.cosmos.accountKey": accountKey,
    "spark.cosmos.database": databaseName,
    "spark.cosmos.container": containerName,
    "spark.cosmos.read.customQuery": "SELECT * FROM c WHERE c.type = '你需要的文档类型值'"
}

df = spark.read.format("cosmos.oltp").options(**cfg).schema(target_schema).load()

内容的提问来源于stack exchange,提问作者jukebox

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 01:24:01