Spark中将字符串转为DoubleType后字段全为Null的技术问题
Spark转换CSV带逗号的数值字段为Double类型时全Null问题解决
问题背景
- 数据集的CSV片段:
date,% Iron Feed,% Silica Feed 2017-03-10 01:00:00,"55,2","16,98" 2017-03-10 01:00:00,"55,2","16,98" 2017-03-10 01:00:00,"55,2","16,98"
- 执行的Spark代码:
sourcefile = 'MiningProcess_Flotation_Plant_Database.csv' df = spark.read.format('csv').option("header","true").load(db_ws.dp_engagement + '/' + sourcefile) display(df) from pyspark.sql.functions import col from pyspark.sql.types import StringType,BooleanType,DateType,DoubleType df2 = df.withColumn("% Iron Feed",col("% Iron Feed").cast(DoubleType())) df2.printSchema()
- 结果:Schema显示
% Iron Feed已设为double类型,但该字段所有值均为Null。
问题原因
CSV中的数值使用逗号作为小数分隔符(例如55,2),而Spark默认以点号作为小数分隔符,直接执行cast(DoubleType())会因格式不匹配导致转换失败,最终生成Null值。
解决方案
方案1:读取CSV时指定小数分隔符
读取阶段直接配置小数分隔符,让Spark自动识别数值类型,无需后续转换:
sourcefile = 'MiningProcess_Flotation_Plant_Database.csv' df = spark.read.format('csv')\ .option("header", "true")\ .option("decimalSeparator", ",")\ .load(db_ws.dp_engagement + '/' + sourcefile) display(df) df.printSchema()
方案2:替换逗号为点号后再转换
如果已经完成数据读取,可通过字符串替换修正格式后再转换类型:
from pyspark.sql.functions import col, regexp_replace # 处理% Iron Feed字段 df2 = df.withColumn("% Iron Feed", regexp_replace(col("% Iron Feed"), ",", ".").cast(DoubleType())) # 同理处理% Silica Feed字段 df2 = df2.withColumn("% Silica Feed", regexp_replace(col("% Silica Feed"), ",", ".").cast(DoubleType())) display(df2) df2.printSchema()
内容的提问来源于stack exchange,提问作者Luis Valencia
相关产品推荐
相关产品推荐

