You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何配置PySpark读取CSV时将空单元格识别为空字符串?

解决PySpark读取CSV时空单元格识别为空字符串的问题

有两种实用方法可以实现需求:

方法一:读取CSV时配置参数直接解析为空字符串

在spark.read.csv中通过参数配置,让空单元格直接被解析为空字符串,避免被识别为Null:

from pyspark.sql.types import StructType, StructField, StringType

schema = StructType([
    StructField("col1", StringType(), nullable=False),
    StructField("col2", StringType(), nullable=False),
    StructField("col3", StringType(), nullable=False)
])
file1 = 'file.csv'
# 配置空单元格对应空字符串,用自定义特殊值标记真正的Null(避免和空单元格混淆)
df1 = spark.read.csv(
    file1,
    header=True,
    schema=schema,
    emptyValue="",
    nullValue="__CUSTOM_NULL__"
)
df1.show()

这种方式会直接把CSV里的空单元格解析为空字符串,完全匹配你要的输出效果,同时因为提前规避了Null,也符合你schema中nullable=False的设置。

方法二:读取后批量替换Null为空字符串

如果已经完成CSV读取,可通过fillna方法快速将DataFrame中所有Null值替换为空字符串:

from pyspark.sql.types import StructType, StructField, StringType

schema = StructType([
    StructField("col1", StringType(), nullable=False),
    StructField("col2", StringType(), nullable=False),
    StructField("col3", StringType(), nullable=False)
])
file1 = 'file.csv'
df1 = spark.read.csv(file1, header=True, schema=schema)
# 批量替换所有列的Null为空字符串
df1 = df1.fillna("")
df1.show()

注意:由于你的schema设置了nullable=False,读取时若出现Null会触发报错,因此更推荐使用方法一提前处理。


内容的提问来源于stack exchange,提问作者Rakesh Kushwaha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 19:08:34