You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark 3.3.1中CORRECTED模式下仍触发timeParserPolicy升级异常

Spark 3.3.1日期解析异常问题

环境信息

  • Spark版本:3.3.1
  • Python版本:3.9

问题背景

默认情况下,当foobar列的日期字符串无法通过指定格式(如yyyy-MM-dd)解析,但旧版日期解析器可以解析时,调用pyspark.sql.functions.to_date(col("foobar"), "yyyy-MM-dd")会触发Spark升级异常。

异常信息

org.apache.spark.SparkUpgradeException: 
You may get a different result due to the upgrading of Spark 3.0: 
Fail to parse '21/09/20' in the new parser.

You can set spark.sql.legacy.timeParserPolicy to
LEGACY to restore the behavior before Spark 3.0, 
or set to CORRECTED and treat it as an invalid datetime string.

矛盾点

Spark会话已配置为CORRECTED模式:

spark.conf.get("spark.sql.legacy.timeParserPolicy") # 返回CORRECTED

使用CORRECTED模式的预期是保持严格日期校验,解析失败时返回null而非触发异常,便于后续优雅处理,但复杂作业中仍触发异常。

简单示例验证

通过简单代码复现场景,结果符合预期:

  • 未配置CORRECTED时,触发Spark升级异常
  • 配置CORRECTED后,解析失败的记录返回null,无异常

示例代码

spark = SparkSession \
    .builder \
    .appName("foo") \
    .getOrCreate()


dates_df = spark.createDataFrame(
    data=[
        (1, "2024-01-01"),
        (1, "2024-1-01"),
        (1, "2024-01-1"),
        (1, "2024-1-1"),
        (1, "24-01-01"),
        (1, "24-1-01"),
        (1, "24-01-1"),
        (1, "24-1-1"),
        (1, "2024/01/01"),
        (1, "2024/1/01"),
        (1, "2024/01/1"),
        (1, "2024/1/1"),
        (1, "24/01/01"),
        (1, "24/1/01"),
        (1, "24/01/1"),
        (1, "24/1/1"),
        (1, "11/01/2024"),
        (1, "11/1/2024"),
        (1, "11/01/24"),
        (1, "11/1/2024"),
    ],
    schema=["id", "raw_date"],
)
dates_df = dates_df.drop("id")

format = "yyyy-MM-dd"
dates_df.withColumn("date_formatted", to_date("raw_date", format)).show()

未配置CORRECTED时的异常

org.apache.spark.SparkUpgradeException: 
You may get a different result due to the upgrading to Spark >= 3.0: 
Fail to parse '2024-1-01' in the new parser. 
You can set spark.sql.legacy.timeParserPolicy to 
LEGACY to restore the behavior before Spark 3.0, 
or set to CORRECTED and treat it as an invalid datetime string.

配置CORRECTED后的结果

+----------+--------------+
|  raw_date|date_formatted|
+----------+--------------+
|2024-01-01|    2024-01-01|
| 2024-1-01|          null|
| 2024-01-1|          null|
|  2024-1-1|          null|
|  24-01-01|          null|
|   24-1-01|          null|
|   24-01-1|          null|
|    24-1-1|          null|
|2024/01/01|          null|
| 2024/1/01|          null|
| 2024/01/1|          null|
|  2024/1/1|          null|
|  24/01/01|          null|
|   24/1/01|          null|
|   24/01/1|          null|
|    24/1/1|          null|
|11/01/2024|          null|
| 11/1/2024|          null|
|  11/01/24|          null|
| 11/1/2024|          null|
+----------+--------------+

复杂作业中的异常表现

该问题仅在复杂/长作业中出现,且不同环境下触发位置不同:

  • 本地单节点环境:作业初期(第2/10阶段)触发异常
  • Kubernetes集群远程环境:作业末尾执行df.rdd.getNumPartitions()时触发异常(本地运行该代码段无问题)

尝试过的解决方法

推测问题与资源、分区及作业复杂度有关,尝试通过checkpoint截断逻辑计划,但未解决问题:

with tempfile.TemporaryDirectory() as d:
    spark.sparkContext.setCheckpointDir("/tmp/bb")
    df = df.checkpoint(True)

即使配置了CORRECTED模式,仍会触发提示切换至LEGACY或CORRECTED的升级异常。

需求说明

作业中仅使用to_date(col('foobar'), '<date_format>')进行日期解析,且确实运行在CORRECTED模式下。LEGACY模式虽能正常执行作业,但日期校验不够严格,无法满足业务需求。


内容的提问来源于stack exchange,提问作者hdw3

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 20:14:52