PySpark文本数据源不支持布尔类型报错问题咨询
问题解决:Spark Text数据源不支持Boolean类型报错
问题原因
Spark的text数据源仅支持将每行数据读取为单个字符串列(默认列名为value),不支持自定义包含多列及Boolean等非字符串类型的Schema。你试图用text格式加载带Boolean类型的结构化分隔数据,这就触发了报错。
解决方案
你的数据是空格分隔的结构化数据,应改用支持解析多列的数据源格式(如csv),同时处理数据中的异常值并转换类型:
步骤1:定义Schema并以CSV格式加载数据
首先导入所需类型,定义包含Integer类型的Schema(先避开Boolean,后续转换),然后用csv格式加载并指定空格分隔符:
from pyspark.sql.types import StructType, StructField, StringType, IntegerType, BooleanType # 定义Schema:present先设为IntegerType,后续转Boolean schema = StructType([ StructField("name", StringType(), nullable=True), StructField("present", IntegerType(), nullable=True), StructField("country", StringType(), nullable=True) ]) # 用CSV格式加载,指定多空格分隔符、读取表头 df = spark.read.format("csv")\ .option("sep", r"\s+")\ .option("header", "true")\ .schema(schema)\ .load("a.txt")
步骤2:处理异常值并转换为Boolean类型
数据中存在1.这种带小数点的数值,先清理后再转换为Boolean:
from pyspark.sql.functions import regexp_replace # 去掉present字段的小数点,再转为Integer,最后转Boolean df = df.withColumn("present", regexp_replace(df["present"], r"\.", "").cast(IntegerType()))\ .withColumn("present", df["present"].cast(BooleanType()))
备选方案:先读取文本再手动拆分列
如果必须用text格式读取,可先读取所有行,再手动拆分列并转换类型:
from pyspark.sql.functions import split # 读取所有文本行 df_text = spark.read.text("a.txt") # 获取表头并拆分列名 header_row = df_text.first().value columns = header_row.split() # 过滤表头行,拆分数据列 df_data = df_text.filter(df_text.value != header_row)\ .select(*[split(df_text.value, r"\s+")[i].alias(col) for i, col in enumerate(columns)]) # 转换present字段类型 df_data = df_data.withColumn("present", regexp_replace(df_data["present"], r"\.", "").cast(IntegerType()).cast(BooleanType()))
内容的提问来源于stack exchange,提问作者Xi12
相关产品推荐
相关产品推荐

