You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Pandas DataFrame通过循环构建嵌套StructType结构

基于Pandas DataFrame循环构建嵌套StructType结构

给定以下Pandas数据结构:

import pandas as pd
import numpy as np

a1=["data.country", "data.studentinfo.city","data.studentinfo.name.id.grant"]
a2=["StringType()","StringType()","StringType()"]
d1=pd.DataFrame(list(zip(a1,a2)),columns=['action','type'])

需求是通过for循环,基于上述DataFrame构建出嵌套的Spark StructType结构,最终目标结构如下:

StructType([StructField("data", 
    StructType([StructField("country",StringType(),True),
                StructField("studentinfo",
                StructType([StructField("city",StringType(),True),
                    StructField("name",StructType([
                        StructField("id",StructType([
                        StructField("grant",StringType(),True)])
                        ]))
                ])
            )])
    )])

解决方案

首先导入Spark所需的类型模块:

from pyspark.sql.types import StructType, StructField, StringType

接着编写循环构建嵌套结构的代码:

# 初始化根StructType
root_struct = StructType()
# 用字典跟踪各层级的StructType实例,便于快速定位父结构
struct_track = {"": root_struct}

for _, row in d1.iterrows():
    # 拆分字段的层级路径和对应的类型
    path_parts = row['action'].split('.')
    # 安全起见,建议用类型映射替代eval,示例见下方说明
    field_type = eval(row['type'])
    
    current_path = ""
    # 遍历路径的每一层(除最后一个字段名),创建缺失的层级StructType
    for part in path_parts[:-1]:
        full_path = f"{current_path}.{part}" if current_path else part
        if full_path not in struct_track:
            # 获取父结构,创建新的子StructType并添加进去
            parent_struct = struct_track[current_path]
            child_struct = StructType()
            parent_struct.add(StructField(part, child_struct, True))
            struct_track[full_path] = child_struct
        current_path = full_path
    
    # 将最底层字段添加到对应的父结构中
    parent_struct = struct_track[current_path]
    parent_struct.add(StructField(path_parts[-1], field_type, True))

# 最终生成的嵌套StructType
final_struct = root_struct

注意事项

代码中使用eval(row['type'])仅适用于完全信任输入数据的场景,若输入存在不可信内容,建议改用类型映射字典来避免安全风险:

# 定义类型映射表
type_map = {
    "StringType()": StringType(),
    # 可根据需求添加其他Spark数据类型,比如IntegerType()、FloatType()等
}

# 替换原代码中的field_type赋值行
field_type = type_map[row['type']]

内容的提问来源于stack exchange,提问作者Amol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 04:46:47