PySpark读取百万级JSON文件遇Unable to infer schema错误的解决方法
解决PySpark读取大JSON文件无法推断Schema的问题
方法1:手动指定Schema
既然PySpark自动推断schema失败,直接用pandas先拿到准确的schema再传给PySpark就行:
- 用pandas读取少量样本数据(比如前1000条),提取字段和类型信息
- 把pandas的数据类型转换成PySpark对应的类型
- 读取文件时手动指定这个schema
示例代码:
import pandas as pd from pyspark.sql.types import StructType, StructField, StringType, IntegerType, FloatType, TimestampType # 用pandas读少量数据拿schema pd_sample = pd.read_json("file.json", encoding='utf-8-sig', nrows=1000) # 定义pandas到PySpark的类型映射,按需补充 dtype_map = { 'object': StringType(), 'int64': IntegerType(), 'float64': FloatType(), 'datetime64[ns]': TimestampType() } # 构建PySpark的StructType spark_schema = StructType([ StructField(col_name, dtype_map[str(pd_sample.dtypes[col_name])], nullable=True) for col_name in pd_sample.columns ]) # 带着schema读取大文件 spark_df = spark.read.option("multiline", "true") \ .option("encoding", "utf-8-sig") \ .schema(spark_schema) \ .json("file.json")
方法2:提高PySpark的采样比例
PySpark默认只采样部分数据推断schema,大文件可能采样到的片段没覆盖全所有字段/类型,直接设置采样全部数据试试:
spark_df = spark.read.option("multiline", "true") \ .option("encoding", "utf-8-sig") \ .option("samplingRatio", "1.0") \ # 采样100%数据 .json("file.json")
注意:文件特别大的话,这个操作会耗时,适合数据格式稳定的场景。
方法3:统一编码设置
你用pandas时指定了utf-8-sig编码,PySpark默认用utf-8,可能因为BOM头问题导致schema推断失败,给PySpark加上编码参数:
spark_df = spark.read.option("multiline", "true") \ .option("encoding", "utf-8-sig") \ .json("file.json")
方法4:预处理成单行JSON格式
multiline模式下PySpark处理大文件容易出问题,先用pandas把文件转成每行一个JSON对象的格式,再用PySpark读取:
# pandas读全量数据 pd_full = pd.read_json("file.json", encoding='utf-8-sig') # 导出为每行一条JSON的格式 pd_full.to_json("formatted_file.json", orient="records", lines=True, force_ascii=False) # PySpark直接读取预处理后的文件,不需要multiline参数 spark_df = spark.read.json("formatted_file.json")
内容的提问来源于stack exchange,提问作者Tavakoli
相关产品推荐
相关产品推荐

