PySpark如何获取array及struct类型列中的嵌套字段名称
PySpark获取DataFrame嵌套字段的原生方法
你之前尝试用df.列名.columns无法拿到嵌套字段,是因为df.列名返回的是PySpark的Column对象,不包含schema元信息,我们可以直接通过解析DataFrame原生的schema属性获取所有层级的字段,不需要引入任何第三方依赖。
1. 单独获取指定嵌套列的字段
针对Array类型列(比如你的array_1)
array_1) # 先获取array_1字段的schema定义 array_col_schema = df.schema["array_1"] # 提取array元素的struct类型的字段名列表 inner_fields = array_col_schema.dataType.elementType.names print(inner_fields) # 输出: ['id_2', 'post']
针对普通Struct类型列
# 假设根层级有struct类型列`struct_1`,获取其内部字段 struct_col_schema = df.schema["struct_1"] inner_fields = struct_col_schema.dataType.names
2. 通用递归函数(获取全量嵌套字段路径)
如果需要批量处理任意层级的嵌套字段,可直接使用下面的递归函数,支持输出所有嵌套字段的完整路径,方便后续循环调用:
from pyspark.sql.types import StructType, ArrayType def get_all_nested_fields(schema, prefix=""): all_fields = [] for field in schema.fields: current_path = f"{prefix}.{field.name}" if prefix else field.name if isinstance(field.dataType, StructType): # 递归处理struct类型的内部字段 all_fields.extend(get_all_nested_fields(field.dataType, current_path)) elif isinstance(field.dataType, ArrayType) and isinstance(field.dataType.elementType, StructType): # 递归处理array内struct的内部字段 all_fields.extend(get_all_nested_fields(field.dataType.elementType, current_path)) else: # 基础数据类型字段直接添加路径 all_fields.append(current_path) return all_fields # 调用示例 full_fields = get_all_nested_fields(df.schema) print(full_fields) # 对应你的示例schema输出为:['id_1', 'array_1.id_2', 'array_1.post.value']
如果你只需要获取某一列下的所有嵌套字段,只需传入对应列的schema即可,比如单独获取array_1下的所有字段:
array_1_all_fields = get_all_nested_fields(df.schema["array_1"].dataType.elementType) print(array_1_all_fields) # 输出: ['id_2', 'post.value']
内容的提问来源于stack exchange,提问作者emilk
相关产品推荐
相关产品推荐

