PySpark如何从混合类型Map提取Struct字段并拆分列?
解决PySpark中Map字段混合类型提取报错问题
报错原因
你遇到的AnalysisException是因为info字段是Map类型,其中address对应的value存在Struct和BigInt两种混合类型。Spark在解析时无法确定info.address的统一类型,直接用.访问Struct字段就会抛出类型不匹配的错误。
解决方案
需要先处理info.address的类型不一致问题,只对Struct类型的记录提取字段,非Struct类型的可设为null或自定义默认值。以下是两种可行的实现方式:
方法1:通过类型判断提取字段
使用typeof函数判断info.address的类型,仅当类型为struct时提取city和state:
from pyspark.sql.functions import col, when, typeof # 模拟包含混合类型的实际数据 data = [ ("Alice", {"age": 25, "address": {"city": "New York", "state": "NY"}}), ("Bob", {"age": 30, "address": {"city": "San Francisco", "state": "CA"}}), ("Charlie", {"age": 35, "address": {"city": "Chicago", "state": "IL"}}), ("Dave", {"age": 40, "address": 123}) # BigInt类型的address ] df = spark.createDataFrame(data, ["name", "info"]) df.select( col("name"), when(typeof(col("info.address")) == "struct", col("info.address.city")).alias("city"), when(typeof(col("info.address")) == "struct", col("info.address.state")).alias("state") ).show()
方法2:强制类型转换提取
将info.address强制转换为指定结构的Struct类型,转换失败的记录对应字段会返回null:
from pyspark.sql.functions import col df.select( col("name"), # 根据实际字段类型定义Struct结构 col("info.address").cast("struct<city:string,state:string>").city.alias("city"), col("info.address").cast("struct<city:string,state:string>").state.alias("state") ).show()
执行结果
两种方法都会输出符合预期的结果(混合类型记录的对应字段为null):
+--------+-------------+-----+ | name| city|state| +--------+-------------+-----+ | Alice| New York| NY| | Bob|San Francisco| CA| | Charlie| Chicago| IL| | Dave| null| null| +--------+-------------+-----+
内容的提问来源于stack exchange,提问作者Sachin Sukumaran
相关产品推荐
相关产品推荐

