Azure Databricks中从CDM JSON生成CSV读取用StructType的方法
解决方案
你可以直接通过PySpark内置的数组展开函数提取元数据,无需用原生Python读取JSON文件,完整步骤如下:
步骤1:展开嵌套数组提取字段元数据
你已经把CDM JSON加载为DataFrame,直接用explode函数两次展开嵌套的数组即可拿到所有字段的name和dataFormat:
from pyspark.sql.functions import explode # 先展开definitions数组,再展开hasAttributes数组 attr_df = df1.select(explode("definitions").alias("def")) \ .select(explode("def.hasAttributes").alias("attr")) \ .select("attr.name", "attr.dataFormat") # 可以先打印验证拿到的字段信息 attr_df.show(truncate=False)
步骤2:做CDM类型到Spark类型的映射
CDM的dataFormat和Spark的字段类型需要做对应,你可以根据实际业务用到的类型补全映射规则:
from pyspark.sql.types import * # 示例映射规则,可按需扩展 type_mapping = { "string": StringType(), "int32": IntegerType(), "int64": LongType(), "decimal": DecimalType(18,2), # 可按需调整精度 "date": DateType(), "datetime": TimestampType(), "boolean": BooleanType(), "double": DoubleType() }
步骤3:构造StructType schema
把提取到的字段信息遍历生成StructField,拼成最终的schema:
# 收集所有字段信息到Python列表 fields = attr_df.collect() # 生成StructType csv_schema = StructType([ StructField(field.name, type_mapping.get(field.dataFormat, StringType()), nullable=True) for field in fields ]) # 打印验证生成的schema print(csv_schema.simpleString())
步骤4:用生成的schema读取无表头CSV
csv_df = spark.read.schema(csv_schema) \ .option("header", "false") \ .csv("/mnt/你的CSV文件路径")
补充说明
如果确实需要用原生Python读取该JSON文件,只需要给路径加上/dbfs前缀即可,因为Databricks的/mnt/挂载点对应本地文件系统的/dbfs/mnt/路径:
import json with open("/dbfs/mnt/jsontest/...PATH.../SalesTable.cdm.json", "r", encoding="utf-8") as f: cdm_json = json.load(f) # 后续直接解析cdm_json字典即可
内容的提问来源于stack exchange,提问作者user1567438
相关产品推荐
相关产品推荐

