You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark如何获取array及struct类型列中的嵌套字段名称

PySpark获取DataFrame嵌套字段的原生方法

你之前尝试用df.列名.columns无法拿到嵌套字段,是因为df.列名返回的是PySpark的Column对象,不包含schema元信息,我们可以直接通过解析DataFrame原生的schema属性获取所有层级的字段,不需要引入任何第三方依赖。

1. 单独获取指定嵌套列的字段

针对Array类型列(比如你的array_1)

# 先获取array_1字段的schema定义
array_col_schema = df.schema["array_1"]
# 提取array元素的struct类型的字段名列表
inner_fields = array_col_schema.dataType.elementType.names
print(inner_fields)
# 输出: ['id_2', 'post']

针对普通Struct类型列

# 假设根层级有struct类型列`struct_1`,获取其内部字段
struct_col_schema = df.schema["struct_1"]
inner_fields = struct_col_schema.dataType.names

2. 通用递归函数(获取全量嵌套字段路径)

如果需要批量处理任意层级的嵌套字段,可直接使用下面的递归函数,支持输出所有嵌套字段的完整路径,方便后续循环调用:

from pyspark.sql.types import StructType, ArrayType

def get_all_nested_fields(schema, prefix=""):
    all_fields = []
    for field in schema.fields:
        current_path = f"{prefix}.{field.name}" if prefix else field.name
        if isinstance(field.dataType, StructType):
            # 递归处理struct类型的内部字段
            all_fields.extend(get_all_nested_fields(field.dataType, current_path))
        elif isinstance(field.dataType, ArrayType) and isinstance(field.dataType.elementType, StructType):
            # 递归处理array内struct的内部字段
            all_fields.extend(get_all_nested_fields(field.dataType.elementType, current_path))
        else:
            # 基础数据类型字段直接添加路径
            all_fields.append(current_path)
    return all_fields

# 调用示例
full_fields = get_all_nested_fields(df.schema)
print(full_fields)
# 对应你的示例schema输出为:['id_1', 'array_1.id_2', 'array_1.post.value']

如果你只需要获取某一列下的所有嵌套字段,只需传入对应列的schema即可,比如单独获取array_1下的所有字段:

array_1_all_fields = get_all_nested_fields(df.schema["array_1"].dataType.elementType)
print(array_1_all_fields)
# 输出: ['id_2', 'post.value']

内容的提问来源于stack exchange,提问作者emilk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 22:57:03