PySpark中使用createDataFrame转换含None的列表为DataFrame报错求解
解决包含Null值的列表创建PySpark DataFrame的问题
问题原因
当列表中某一列只有None值时,PySpark无法推断该列的数据类型,因此抛出ValueError: Some of types cannot be determined after inferring错误。
解决方案:显式指定Schema
通过StructType和StructField明确定义每一列的数据类型,让PySpark无需自行推断,同时允许列值为Null。
步骤1:导入所需类型
from pyspark.sql.types import StructType, StructField, StringType # 若涉及整数、时间戳等其他类型,需对应导入IntegerType、TimestampType等
步骤2:定义Schema结构
根据数据实际类型调整字段类型(示例假设ID、状态、时间字段均为字符串类型):
schema = StructType([ StructField("RUN_ID", StringType(), nullable=True), StructField("RE_RUN_ID", StringType(), nullable=True), StructField("CONFIG_ID", StringType(), nullable=True), StructField("JOB_STATUS", StringType(), nullable=True), StructField("LOCK_STATUS", StringType(), nullable=True), StructField("INSERT_DTS", StringType(), nullable=True), StructField("UPDATE_DTS", StringType(), nullable=True) ])
步骤3:创建DataFrame
将定义好的Schema传入createDataFrame方法:
row = [Row(run_id, None, self.args.config_id, 'in-progress', 'locked', datetime.now().strftime('%Y-%m-%d %H:%M:%S'), datetime.now().strftime('%Y-%m-%d %H:%M:%S') ) ] df = sc.spark.createDataFrame(row, schema=schema) df.show(truncate=False)
验证Null值
执行验证代码,此时flag列会返回1,说明None已被正确识别为PySpark的Null值:
df.withColumn("flag", fn.when(fn.col("RE_RUN_ID").isNull(), 1).otherwise(0)).show()
内容的提问来源于stack exchange,提问作者shankar
相关产品推荐
相关产品推荐

