如何在Spark Scala 3.2.1中创建包含None/Null值的DataFrame
Spark Scala 3.2.1 初始化含None/Null值的DataFrame解决方法
错误原因
你写的代码里,category2列同时存在Boolean类型(true/false)和None,导致该列的推断类型变成了Any,而Spark不支持将Any作为DataFrame的列类型,因此抛出UnsupportedOperationException: Schema for type Any is not supported错误。
解决方法
方法1:统一使用Option类型包装所有值
将所有非Null值用Some()包装,让整列的类型统一为Option[Boolean],Spark会自动推断出允许为空的Boolean类型Schema:
val df = spark.createDataFrame( Seq( (0, "a", Some(true)), (1, "b", Some(true)), (2, "c", Some(false)), (3, "a", Some(false)), (4, "a", None), (5, "c", Some(false)) ) ).toDF("id", "category1", "category2") df.show() df.printSchema()
执行后category2的Schema会显示为boolean (nullable = true),符合需求。
方法2:手动定义Schema并使用Row构造数据
如果需要更精确控制Schema,可以手动定义StructType,然后用Row来构造数据,Null值用null表示:
import org.apache.spark.sql.types.{StructType, StructField, IntegerType, StringType, BooleanType} import org.apache.spark.sql.Row val schema = StructType(Seq( StructField("id", IntegerType, nullable = false), StructField("category1", StringType, nullable = false), StructField("category2", BooleanType, nullable = true) )) val df = spark.createDataFrame( spark.sparkContext.parallelize(Seq( Row(0, "a", true), Row(1, "b", true), Row(2, "c", false), Row(3, "a", false), Row(4, "a", null), Row(5, "c", false) )), schema ) df.show() df.printSchema()
这种方式直接指定了每个列的可空性,适合复杂Schema场景。
内容的提问来源于stack exchange,提问作者yanachen
相关产品推荐
相关产品推荐

