PySpark展开JSON列遇schema问题:'tuple'对象无'name'属性如何解决
错误原因排查
- 核心错误来自schema定义的语法问题:你给
StructType传入的列表中,用括号包裹了两个StructField对象,形成了嵌套元组,StructType仅能接收由StructField单独构成的列表,无法识别元组对象,因此抛出'tuple' object has no attribute 'name'报错。 - 次要问题:代码中直接使用
col()函数但未导入该方法,后续运行也会触发报错。 - 额外注意:你已经通过
spark.read.json读取了GeoJSON文件,如果读取时没有强制指定schema为字符串,geometry列本身已经是Struct类型,不需要再调用from_json解析,可直接提取子字段。
修复方案
方案1:若geometry列已经是Struct类型(绝大多数场景)
直接提取子字段即可,不需要额外解析:
from pyspark.sql.types import StructField, StructType, StringType, FloatType, ArrayType, DoubleType import pyspark.sql.functions as F df = spark.read.option("multiLine", False).option("mode", "PERMISSIVE").json('Italy/it_countrywide-addresses-country.geojson') # 直接提取geometry的子字段 df.select(F.col("geometry.*")).show()
方案2:若geometry列确实为字符串类型需要解析
修正schema定义语法即可:
from pyspark.sql.types import StructField, StructType, StringType, FloatType, ArrayType, DoubleType import pyspark.sql.functions as F df = spark.read.option("multiLine", False).option("mode", "PERMISSIVE").json('Italy/it_countrywide-addresses-country.geojson') # 修正schema定义,去掉多余的括号,直接传StructField列表 geometry_schema = StructType([ StructField("coordinates", ArrayType(DoubleType()), True), StructField("type", StringType(), True) ]) df.withColumn("geometry", F.from_json("geometry", geometry_schema))\ .select(F.col('geometry.*'))\ .show()
内容的提问来源于stack exchange,提问作者Tim
相关产品推荐
相关产品推荐

