PySpark读取CSV指定DateType schema仍解析为StringType问题排查
问题1:指定DateType后字段仍为StringType的原因
核心原因是schema传参方式错误:
- Spark的DataFrameReader中,schema是独立的配置入口,不属于普通option选项,你使用
.option("schema", school_students_schema)的写法,Spark会直接静默忽略这个无效配置,既不会应用你定义的StructType,也不会抛出任何异常。 - 因为自定义schema没有生效,读取器默认又没有开启
inferSchema自动类型推断,会将所有CSV字段默认解析为StringType,你定义的doj字段类型、配置的dateFormat规则都没有被实际应用到解析流程中。
问题2:正确解析DateType的修改方法
只需要把schema的传入方式从option配置改为调用DataFrameReader的.schema()方法即可,你配置的dateFormat选项本身写法是正确的,不需要调整。
修正后的可运行代码:
from pyspark.sql.types import StructField, StructType, StringType, DateType school_students_schema = StructType([ StructField("school_id", StringType(), True), StructField("gender", StringType(), True), StructField("class", StringType(), True), StructField("doj", DateType(), True) ]) school_students_df = spark.read.format("csv") \ .option("header", True) \ .schema(school_students_schema) \ .option("dateFormat", "dd/MM/yyyy") \ .load("/user/test/school_students.csv") school_students_df.printSchema()
执行后输出的schema符合预期:
root |-- school_id: string (nullable = true) |-- gender: string (nullable = true) |-- class: string (nullable = true) |-- doj: date (nullable = true)
注:Spark对读写操作中传入的未定义option键默认采取静默丢弃策略,不会抛出报错提示,后续如果遇到自定义配置不生效的场景,可以优先排查是否用错了配置入口。
内容的提问来源于stack exchange,提问作者Monami Sen
相关产品推荐
相关产品推荐

