Spark-XML读取S3文件时显式定义Schema遇到Array类型语法报错问题
问题排查与解决方案
错误根因
StructType构造器仅接收单个StructField数组作为入参,你将BroadcastMetadata、Lines两个字段拆分为两个独立数组传入,不符合语法要求ArrayType入参要求为数组元素的DataType类型,你直接传入StructField触发类型不匹配报错- spark-xml读取重复标签生成的数组结构不需要额外定义
element层,直接声明数组元素的结构即可
正确Schema定义代码
import org.apache.spark.sql.types._ val schema = StructType( Array( // BroadcastMetadata结构定义 StructField("BroadcastMetadata", StructType( Array( StructField("ProgramInfo", StructType( Array( // 样例中该字段为Long类型,可按需保留StringType或改为LongType StructField("_ProgramInfoID", StringType, nullable = true) ) ), nullable = true) ) ), nullable = true), // Lines结构定义 StructField("Lines", StructType( Array( // Line为数组类型,数组元素为包含所需字段的结构体 StructField("Line", ArrayType( StructType( Array( StructField("_VALUE", StringType, nullable = true) // 如需提取Line下的其他字段,直接在当前数组新增StructField即可 ) ), containsNull = true ), nullable = true) ) ), nullable = true) ) )
后续如果需要扩展提取其他字段,只需要在对应层级StructType的数组中新增StructField定义即可,完全匹配你提供的样例Schema结构。
内容的提问来源于stack exchange,提问作者swaythecat
相关产品推荐
相关产品推荐

