PySpark中如何将嵌套数组列FilteredOutDecisions.Reasons转为Double类型
解决方案
要将嵌套在FilteredOutDecisions数组结构体中的Reasons整数数组转换为Double类型数组,可以利用PySpark的transform函数结合withField方法实现嵌套字段的类型转换,具体步骤如下:
1. 导入PySpark函数库
from pyspark.sql import functions as F
2. 执行类型转换
通过嵌套的transform函数遍历数组并修改字段类型:
# 处理FilteredOutDecisions数组中的每个结构体元素 df = df.withColumn( "FilteredOutDecisions", F.transform( "FilteredOutDecisions", # 遍历每个结构体,更新Reasons字段的类型 lambda struct_element: struct_element.withField( "Reasons", # 将Reasons数组中的每个integer元素转换为double F.transform(struct_element["Reasons"], lambda num: num.cast("double")) ) ) )
3. 验证转换结果
执行转换后,可通过打印Schema确认类型是否更新:
df.printSchema()
转换后的Schema中,FilteredOutDecisions.element.Reasons的元素类型会从integer变为double:
root |-- SortedLenders: array (nullable = true) | |-- element: struct (containsNull = true) | | |-- LenderID: string (nullable = true) | | |-- MaxProfit: string (nullable = true) |-- FilteredOutDecisions: array (nullable = true) | |-- element: struct (containsNull = true) | | |-- ApprovedAmount: integer (nullable = true) | | |-- Reasons: array (nullable = true) | | | |-- element: double (containsNull = true)
内容的提问来源于stack exchange,提问作者user12253044
相关产品推荐
相关产品推荐

