You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark从JSON动态创建Schema报错:'str'对象无'name'属性

问题排查:Spark Schema构建报错'str' object has no attribute 'name'

问题背景

在Databricks Notebook中通过Spark从API采集数据,已将全量响应数据存入DataFrame df,需要提取指定列并按定义的类型转换。列名及对应数据类型存储在如下JSON文件中:

{
    "structure": [
        {
            "column_name": "column1",
            "column_type": "StringType()"
        },
        {
            "column_name": "column2",
            "column_type": "IntegerType()"
        },
        {
            "column_name": "column3",
            "column_type": "DateType()"
        },
        {
            "column_name": "column4",
            "column_type": "StringType()"
        }
    ]
}

编写了以下代码构建Schema并创建目标DataFrame:

with open("/dbfs/mnt/datalake/Dims/shema_json","r") as read_handle:
    file_contents = json.load(read_handle)

struct_fields = []
for column in file_contents.get("structure"):
    struct_fields.append(f'StructField(\"{column.get(\"column_name\")}\",{column.get(\"column_type\")},True)')
new_schema = StructType(struct_fields)

df_staging = spark.createDataFrame(df.rdd,schema = new_schema)   

执行时触发错误:'str' object has no attribute 'name'

错误原因

核心问题是代码中将StructField的定义拼接成了字符串,而非创建实际的StructField实例。StructType要求传入的是StructField对象的列表,而不是字符串列表。Spark尝试处理这些字符串时,会将其当作对象访问name属性,自然触发不存在属性的报错。

解决方案

需要将JSON中的类型字符串映射为实际的Spark数据类型对象,再创建StructField实例。修正后的代码如下:

import json
from pyspark.sql.types import StructType, StructField, StringType, IntegerType, DateType

# 读取JSON结构文件
with open("/dbfs/mnt/datalake/Dims/shema_json","r") as read_handle:
    file_contents = json.load(read_handle)

# 建立类型字符串到Spark类型对象的映射
type_mapping = {
    "StringType()": StringType(),
    "IntegerType()": IntegerType(),
    "DateType()": DateType()
}

struct_fields = []
for column in file_contents.get("structure"):
    col_name = column.get("column_name")
    # 从映射中获取对应类型对象
    col_type = type_mapping[column.get("column_type")]
    # 添加StructField实例到列表
    struct_fields.append(StructField(col_name, col_type, nullable=True))

new_schema = StructType(struct_fields)

# 基于正确Schema创建目标DataFrame
df_staging = spark.createDataFrame(df.rdd, schema=new_schema)

优化说明

如果后续JSON中的类型名改为不带括号的格式(如"StringType"),可以用动态获取全局对象的方式简化映射:

type_mapping = {
    "StringType": StringType,
    "IntegerType": IntegerType,
    "DateType": DateType
}
# 循环内创建类型对象
col_type = type_mapping[column.get("column_type")]()

内容的提问来源于stack exchange,提问作者gip

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 17:45:38