如何获取Spark DataFrame中struct类型字段student_info的子列列表?
获取Spark DataFrame中struct类型列的子列列表
针对你给出的DataFrame Schema,要提取student_info下的子列列表,可通过以下方式实现:
Python 实现
假设你的DataFrame变量名为df,执行以下代码:
# 获取student_info字段的Schema信息 struct_field = df.schema["student_info"] # 遍历StructType的字段,提取名称 sub_columns = [field.name for field in struct_field.dataType.fields] print(sub_columns)
执行后会输出:['firstname', 'lastname', 'major', 'hounour_roll']
Scala 实现
假设你的DataFrame变量名为df,执行以下代码:
// 获取student_info字段的Schema信息 val structField = df.schema("student_info") // 转换为StructType并提取子列名称 val subColumns = structField.dataType.asInstanceOf[org.apache.spark.sql.types.StructType].fields.map(_.name) println(subColumns.mkString("[", ", ", "]"))
执行后会输出:[firstname, lastname, major, hounour_roll]
核心逻辑是:先定位到目标struct类型的列,获取其对应的StructType类型对象,再遍历该对象的字段集合,提取每个字段的名称即可。
内容的提问来源于stack exchange,提问作者Pari
相关产品推荐
相关产品推荐

