Spark 3.2.1读取Parquet类型不兼容,如何指定表Schema?
解决Parquet类型不匹配的问题
核心原因
报错明确说明:Parquet文件中ID列的类型是INT32,但你创建表时指定的BIGINT对应Spark的LongType(匹配Parquet的INT64),两者类型不兼容导致读取失败。
解决方案
1. SQL 直接修改表定义
将ID列的类型从BIGINT改为INT(Spark SQL中INT对应Parquet的INT32):
create database if not exists sampledb; drop table if exists sampledb.table; create table sampledb.table (ID INT, Column1 string) using parquet OPTIONS(path='/path/to/parquet/');
2. PySpark 显式指定Schema
通过定义匹配的Schema读取文件并创建表:
from pyspark.sql.types import StructType, StructField, IntegerType, StringType # 定义与Parquet文件匹配的Schema schema = StructType([ StructField("ID", IntegerType(), nullable=True), StructField("Column1", StringType(), nullable=True) ]) # 读取Parquet并创建永久表 df = spark.read.schema(schema).parquet("/path/to/parquet/") df.write.mode("overwrite").saveAsTable("sampledb.table")
3. Scala 显式指定Schema
同样通过Schema匹配实现读取:
import org.apache.spark.sql.types.{StructType, StructField, IntegerType, StringType} // 定义匹配的Schema val schema = StructType(Seq( StructField("ID", IntegerType, nullable = true), StructField("Column1", StringType, nullable = true) )) // 读取文件并生成永久表 spark.read.schema(schema).parquet("/path/to/parquet/") .write.mode("overwrite").saveAsTable("sampledb.table")
内容的提问来源于stack exchange,提问作者Jose R
相关产品推荐
相关产品推荐

