PySpark读取带表头CSV:指定列并强制Schema的实现方法
解决Spark读取CSV时指定Schema并选择特定列的问题
问题原因分析
你的代码未达预期的核心原因:
- 当同时设置
schema和header=True时,Spark会按字段名匹配CSV表头,而非列的位置。你定义的schema包含id和dob,会匹配CSV中表头为id和dob的列,但实际需要的是name和dob列,字段名不匹配导致逻辑错误。 - 错误地将CSV中第二列(字符串类型的
name)映射到schema的dob(DateType),导致日期解析失败,最终返回null。
方案一:按列名匹配(推荐,已知表头时使用)
直接定义与目标列名一致的schema,结合header=True读取,同时指定日期格式确保正确解析:
from pyspark.sql.types import StructType, StructField, StringType, DateType # 定义与目标列名完全匹配的schema schema = StructType([ StructField('name', StringType(), nullable=True), StructField('dob', DateType(), nullable=True) ]) # 读取CSV,指定schema、表头和日期格式 df = spark.read.csv( "somefile.csv", schema=schema, header=True, dateFormat="yyyy-MM-dd" ) df.show()
执行后输出:
+-----+----------+ | name| dob| +-----+----------+ |Smith|2000-01-01| | John|2000-02-01| +-----+----------+
方案二:按位置选择列(无需依赖表头名)
如果需要固定选择第2、3列(不管表头名),可以跳过表头行,按位置读取后重命名列:
from pyspark.sql.types import StructType, StructField, StringType, DateType # 定义对应第2、3列的临时schema schema = StructType([ StructField('temp_name', StringType(), nullable=True), StructField('temp_dob', DateType(), nullable=True) ]) # 读取CSV:跳过表头行,使用schema,指定日期格式 df = spark.read.csv( "somefile.csv", schema=schema, header=False, dateFormat="yyyy-MM-dd", skipRows=1 # 跳过第一行表头 ).withColumnRenamed("temp_name", "name").withColumnRenamed("temp_dob", "dob") df.show()
该方案适用于表头不确定,但列位置固定的场景。
内容的提问来源于stack exchange,提问作者user2535517
相关产品推荐
相关产品推荐

