You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark中将字符串转为DoubleType后字段全为Null的技术问题

Spark转换CSV带逗号的数值字段为Double类型时全Null问题解决

问题背景

  • 数据集的CSV片段:
date,% Iron Feed,% Silica Feed
2017-03-10 01:00:00,"55,2","16,98"
2017-03-10 01:00:00,"55,2","16,98"
2017-03-10 01:00:00,"55,2","16,98"
  • 执行的Spark代码:
sourcefile = 'MiningProcess_Flotation_Plant_Database.csv'
df = spark.read.format('csv').option("header","true").load(db_ws.dp_engagement + '/' + sourcefile)
display(df)
from pyspark.sql.functions import col
from pyspark.sql.types import StringType,BooleanType,DateType,DoubleType
df2 = df.withColumn("% Iron Feed",col("% Iron Feed").cast(DoubleType())) 
df2.printSchema()
  • 结果:Schema显示% Iron Feed已设为double类型,但该字段所有值均为Null。

问题原因

CSV中的数值使用逗号作为小数分隔符(例如55,2),而Spark默认以点号作为小数分隔符,直接执行cast(DoubleType())会因格式不匹配导致转换失败,最终生成Null值。

解决方案

方案1:读取CSV时指定小数分隔符

读取阶段直接配置小数分隔符,让Spark自动识别数值类型,无需后续转换:

sourcefile = 'MiningProcess_Flotation_Plant_Database.csv'
df = spark.read.format('csv')\
    .option("header", "true")\
    .option("decimalSeparator", ",")\
    .load(db_ws.dp_engagement + '/' + sourcefile)
display(df)
df.printSchema()

方案2:替换逗号为点号后再转换

如果已经完成数据读取,可通过字符串替换修正格式后再转换类型:

from pyspark.sql.functions import col, regexp_replace

# 处理% Iron Feed字段
df2 = df.withColumn("% Iron Feed", 
                    regexp_replace(col("% Iron Feed"), ",", ".").cast(DoubleType()))
# 同理处理% Silica Feed字段
df2 = df2.withColumn("% Silica Feed", 
                    regexp_replace(col("% Silica Feed"), ",", ".").cast(DoubleType()))
display(df2)
df2.printSchema()

内容的提问来源于stack exchange,提问作者Luis Valencia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 08:18:24