Spark 3.3.1中CORRECTED模式下仍触发timeParserPolicy升级异常
Spark 3.3.1日期解析异常问题
环境信息
- Spark版本:
3.3.1 - Python版本:
3.9
问题背景
默认情况下,当foobar列的日期字符串无法通过指定格式(如yyyy-MM-dd)解析,但旧版日期解析器可以解析时,调用pyspark.sql.functions.to_date(col("foobar"), "yyyy-MM-dd")会触发Spark升级异常。
异常信息
org.apache.spark.SparkUpgradeException: You may get a different result due to the upgrading of Spark 3.0: Fail to parse '21/09/20' in the new parser. You can set spark.sql.legacy.timeParserPolicy to LEGACY to restore the behavior before Spark 3.0, or set to CORRECTED and treat it as an invalid datetime string.
矛盾点
Spark会话已配置为CORRECTED模式:
spark.conf.get("spark.sql.legacy.timeParserPolicy") # 返回CORRECTED
使用CORRECTED模式的预期是保持严格日期校验,解析失败时返回null而非触发异常,便于后续优雅处理,但复杂作业中仍触发异常。
简单示例验证
通过简单代码复现场景,结果符合预期:
- 未配置
CORRECTED时,触发Spark升级异常 - 配置
CORRECTED后,解析失败的记录返回null,无异常
示例代码
spark = SparkSession \ .builder \ .appName("foo") \ .getOrCreate() dates_df = spark.createDataFrame( data=[ (1, "2024-01-01"), (1, "2024-1-01"), (1, "2024-01-1"), (1, "2024-1-1"), (1, "24-01-01"), (1, "24-1-01"), (1, "24-01-1"), (1, "24-1-1"), (1, "2024/01/01"), (1, "2024/1/01"), (1, "2024/01/1"), (1, "2024/1/1"), (1, "24/01/01"), (1, "24/1/01"), (1, "24/01/1"), (1, "24/1/1"), (1, "11/01/2024"), (1, "11/1/2024"), (1, "11/01/24"), (1, "11/1/2024"), ], schema=["id", "raw_date"], ) dates_df = dates_df.drop("id") format = "yyyy-MM-dd" dates_df.withColumn("date_formatted", to_date("raw_date", format)).show()
未配置CORRECTED时的异常
org.apache.spark.SparkUpgradeException: You may get a different result due to the upgrading to Spark >= 3.0: Fail to parse '2024-1-01' in the new parser. You can set spark.sql.legacy.timeParserPolicy to LEGACY to restore the behavior before Spark 3.0, or set to CORRECTED and treat it as an invalid datetime string.
配置CORRECTED后的结果
+----------+--------------+ | raw_date|date_formatted| +----------+--------------+ |2024-01-01| 2024-01-01| | 2024-1-01| null| | 2024-01-1| null| | 2024-1-1| null| | 24-01-01| null| | 24-1-01| null| | 24-01-1| null| | 24-1-1| null| |2024/01/01| null| | 2024/1/01| null| | 2024/01/1| null| | 2024/1/1| null| | 24/01/01| null| | 24/1/01| null| | 24/01/1| null| | 24/1/1| null| |11/01/2024| null| | 11/1/2024| null| | 11/01/24| null| | 11/1/2024| null| +----------+--------------+
复杂作业中的异常表现
该问题仅在复杂/长作业中出现,且不同环境下触发位置不同:
- 本地单节点环境:作业初期(第2/10阶段)触发异常
- Kubernetes集群远程环境:作业末尾执行
df.rdd.getNumPartitions()时触发异常(本地运行该代码段无问题)
尝试过的解决方法
推测问题与资源、分区及作业复杂度有关,尝试通过checkpoint截断逻辑计划,但未解决问题:
with tempfile.TemporaryDirectory() as d: spark.sparkContext.setCheckpointDir("/tmp/bb") df = df.checkpoint(True)
即使配置了CORRECTED模式,仍会触发提示切换至LEGACY或CORRECTED的升级异常。
需求说明
作业中仅使用to_date(col('foobar'), '<date_format>')进行日期解析,且确实运行在CORRECTED模式下。LEGACY模式虽能正常执行作业,但日期校验不够严格,无法满足业务需求。
内容的提问来源于stack exchange,提问作者hdw3
相关产品推荐
相关产品推荐

