基于Pandas DataFrame通过循环构建嵌套StructType结构
基于Pandas DataFrame循环构建嵌套StructType结构
给定以下Pandas数据结构:
import pandas as pd import numpy as np a1=["data.country", "data.studentinfo.city","data.studentinfo.name.id.grant"] a2=["StringType()","StringType()","StringType()"] d1=pd.DataFrame(list(zip(a1,a2)),columns=['action','type'])
需求是通过for循环,基于上述DataFrame构建出嵌套的Spark StructType结构,最终目标结构如下:
StructType([StructField("data", StructType([StructField("country",StringType(),True), StructField("studentinfo", StructType([StructField("city",StringType(),True), StructField("name",StructType([ StructField("id",StructType([ StructField("grant",StringType(),True)]) ])) ]) )]) )])
解决方案
首先导入Spark所需的类型模块:
from pyspark.sql.types import StructType, StructField, StringType
接着编写循环构建嵌套结构的代码:
# 初始化根StructType root_struct = StructType() # 用字典跟踪各层级的StructType实例,便于快速定位父结构 struct_track = {"": root_struct} for _, row in d1.iterrows(): # 拆分字段的层级路径和对应的类型 path_parts = row['action'].split('.') # 安全起见,建议用类型映射替代eval,示例见下方说明 field_type = eval(row['type']) current_path = "" # 遍历路径的每一层(除最后一个字段名),创建缺失的层级StructType for part in path_parts[:-1]: full_path = f"{current_path}.{part}" if current_path else part if full_path not in struct_track: # 获取父结构,创建新的子StructType并添加进去 parent_struct = struct_track[current_path] child_struct = StructType() parent_struct.add(StructField(part, child_struct, True)) struct_track[full_path] = child_struct current_path = full_path # 将最底层字段添加到对应的父结构中 parent_struct = struct_track[current_path] parent_struct.add(StructField(path_parts[-1], field_type, True)) # 最终生成的嵌套StructType final_struct = root_struct
注意事项
代码中使用eval(row['type'])仅适用于完全信任输入数据的场景,若输入存在不可信内容,建议改用类型映射字典来避免安全风险:
# 定义类型映射表 type_map = { "StringType()": StringType(), # 可根据需求添加其他Spark数据类型,比如IntegerType()、FloatType()等 } # 替换原代码中的field_type赋值行 field_type = type_map[row['type']]
内容的提问来源于stack exchange,提问作者Amol
相关产品推荐
相关产品推荐

