如何配置PySpark读取CSV时将空单元格识别为空字符串?
解决PySpark读取CSV时空单元格识别为空字符串的问题
有两种实用方法可以实现需求:
方法一:读取CSV时配置参数直接解析为空字符串
在spark.read.csv中通过参数配置,让空单元格直接被解析为空字符串,避免被识别为Null:
from pyspark.sql.types import StructType, StructField, StringType schema = StructType([ StructField("col1", StringType(), nullable=False), StructField("col2", StringType(), nullable=False), StructField("col3", StringType(), nullable=False) ]) file1 = 'file.csv' # 配置空单元格对应空字符串,用自定义特殊值标记真正的Null(避免和空单元格混淆) df1 = spark.read.csv( file1, header=True, schema=schema, emptyValue="", nullValue="__CUSTOM_NULL__" ) df1.show()
这种方式会直接把CSV里的空单元格解析为空字符串,完全匹配你要的输出效果,同时因为提前规避了Null,也符合你schema中nullable=False的设置。
方法二:读取后批量替换Null为空字符串
如果已经完成CSV读取,可通过fillna方法快速将DataFrame中所有Null值替换为空字符串:
from pyspark.sql.types import StructType, StructField, StringType schema = StructType([ StructField("col1", StringType(), nullable=False), StructField("col2", StringType(), nullable=False), StructField("col3", StringType(), nullable=False) ]) file1 = 'file.csv' df1 = spark.read.csv(file1, header=True, schema=schema) # 批量替换所有列的Null为空字符串 df1 = df1.fillna("") df1.show()
注意:由于你的schema设置了nullable=False,读取时若出现Null会触发报错,因此更推荐使用方法一提前处理。
内容的提问来源于stack exchange,提问作者Rakesh Kushwaha
相关产品推荐
相关产品推荐

