You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark读取CSV指定DateType schema仍解析为StringType问题排查

问题1:指定DateType后字段仍为StringType的原因

核心原因是schema传参方式错误:

  • Spark的DataFrameReader中,schema是独立的配置入口,不属于普通option选项,你使用.option("schema", school_students_schema)的写法,Spark会直接静默忽略这个无效配置,既不会应用你定义的StructType,也不会抛出任何异常。
  • 因为自定义schema没有生效,读取器默认又没有开启inferSchema自动类型推断,会将所有CSV字段默认解析为StringType,你定义的doj字段类型、配置的dateFormat规则都没有被实际应用到解析流程中。
问题2:正确解析DateType的修改方法

只需要把schema的传入方式从option配置改为调用DataFrameReader的.schema()方法即可,你配置的dateFormat选项本身写法是正确的,不需要调整。
修正后的可运行代码:

from pyspark.sql.types import StructField, StructType, StringType, DateType
school_students_schema = StructType([
    StructField("school_id", StringType(), True),
    StructField("gender", StringType(), True),
    StructField("class", StringType(), True),
    StructField("doj", DateType(), True)    
])

school_students_df = spark.read.format("csv") \
                           .option("header", True) \
                           .schema(school_students_schema) \
                           .option("dateFormat", "dd/MM/yyyy") \
                           .load("/user/test/school_students.csv")
school_students_df.printSchema()

执行后输出的schema符合预期:

root
|-- school_id: string (nullable = true)
|-- gender: string (nullable = true)
|-- class: string (nullable = true)
|-- doj: date (nullable = true)

注:Spark对读写操作中传入的未定义option键默认采取静默丢弃策略,不会抛出报错提示,后续如果遇到自定义配置不生效的场景,可以优先排查是否用错了配置入口。

内容的提问来源于stack exchange,提问作者Monami Sen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 08:03:30