You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark写入分区Parquet后Timestamp字段读取为Null求助

问题:Spark写入分区Parquet后Timestamp字段读取全为Null

问题场景

读取CSV文件后数据和Schema正常,写入按col1分区的Parquet文件后,读取发现col6(Timestamp类型)值全部为Null,但字段类型仍保持Timestamp。

读取CSV后的状态

>>> df.printSchema()
root
 |-- col1: string (nullable = true)
 |-- col5: double (nullable = true)
 |-- col6: timestamp (nullable = true)
 |-- col7: string (nullable = true)

>>> df.show()
+----+-----+-------------------+-------------+
|col1| col5|               col6|         col7|
+----+-----+-------------------+-------------+
|   f| 3.34|1970-01-01 00:00:00|this is test3|
|   f| 2.13|1980-02-05 00:00:00|this is test3|
|   f|12.13|1981-02-05 00:00:00|this is test3|
|   e|  2.3|1982-03-05 00:00:00|this is test3|
|   e|  2.3|1983-04-12 00:00:00|this is test3|
|   e|212.0|1984-05-04 00:00:00|this is test3|
|   e| 2.13|1985-01-10 00:00:00|this is test3|
+----+-----+-------------------+-------------+

写入并读取Parquet后的状态

>>> df.write.partitionBy("col1").mode("append").parquet("<some_location>/testparquetData/")
>>> df1 = spark.read.parquet("<some_location>/testparquetData/")   

>>> df1.show()
+-----+----+-------------+----+
| col5|col6|         col7|col1|
+-----+----+-------------+----+
|  2.3|null|this is test3|   e|
|  2.3|null|this is test3|   e|
|212.0|null|this is test3|   e|
| 2.13|null|this is test3|   e|
| 3.34|null|this is test3|   f|
| 2.13|null|this is test3|   f|
|12.13|null|this is test3|   f|
+-----+----+-------------+----+

>>> df1.printSchema()
root
 |-- col5: double (nullable = true)
 |-- col6: timestamp (nullable = true)
 |-- col7: string (nullable = true)
 |-- col1: string (nullable = true)

已尝试显式指定Schema读取CSV,问题仍存在。


可能原因及解决方案

1. 时区配置不一致

Spark的时区配置会影响Timestamp的存储和解析,若写入和读取时时区不匹配,可能导致Timestamp值解析为Null。

解决办法:在写入和读取前统一设置时区,例如:

# 设置为UTC时区,或根据实际业务时区调整(如"Asia/Shanghai")
spark.conf.set("spark.sql.session.timeZone", "UTC")

2. Parquet Timestamp转换参数未正确配置

Spark对Parquet的INT96格式Timestamp处理需要特定参数,若参数设置错误会导致解析失败。

解决办法:启用INT96格式Timestamp转换(适用于Spark 2.x及以上版本):

spark.conf.set("spark.sql.parquet.int96TimestampConversion", "true")

3. 旧版Parquet格式兼容性问题

若使用了旧版Parquet写入格式,可能导致Timestamp读取异常。

解决办法:禁用旧版写入格式:

spark.conf.set("spark.sql.parquet.writeLegacyFormat", "false")

验证步骤

  1. 执行上述配置后,重新读取CSV并写入Parquet;
  2. 读取新生成的Parquet文件,检查col6字段值是否正常。

内容的提问来源于stack exchange,提问作者Kaushik Ghosh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 07:21:34