Spark读取含非标准格式字段的CSV时如何定义Schema?
Great question! When dealing with non-standard CSV formats like comma decimal separators or dd.MM.yyyy dates, you can avoid reading everything as strings first—Spark has built-in options to handle these cases directly during the read phase. Let's break down your options clearly:
1. Best Approach: Configure CSV Reader Options (No Post-Read Conversion Needed)
Spark's CSV reader supports specific options to adapt to non-standard formats, which lets you define your target schema directly without extra conversion steps. This is more efficient because parsing happens during the initial read.
Key Options to Use:
decimalSeparator: Sets the character used for decimal points (replace the default.with,for your case).dateFormat: Specifies the date pattern to match yourdd.MM.yyyystrings (usetimestampFormatif you're dealing with datetime values instead of just dates).
Java Code Example:
import org.apache.spark.sql.types.*; // Define your target schema with the correct data types StructType customSchema = new StructType() .add("column1", IntegerType, true) // Example integer column .add("double_column", DoubleType, true) // Your comma-separated decimal column .add("date_column", DateType, true); // Your dd.MM.yyyy date column Dataset<Row> ds = sqlContext.read() .format("csv") .option("header", "false") .option("decimalSeparator", ",") // Handle comma as decimal separator .option("dateFormat", "dd.MM.yyyy") // Parse dates in dd.MM.yyyy format .schema(customSchema) .load("myCsvFilePath.csv");
2. Fallback: Read as Strings First, Then Convert
If you run into edge cases (e.g., mixed formats in a single column that the CSV reader can't handle), you can read the problematic columns as strings first, then convert them to the desired type using Spark functions.
Java Code Example:
import static org.apache.spark.sql.functions.*; import org.apache.spark.sql.types.*; // First define a schema with string types for non-standard columns StructType stringSchema = new StructType() .add("column1", StringType, true) .add("double_str", StringType, true) .add("date_str", StringType, true); Dataset<Row> ds = sqlContext.read() .format("csv") .option("header", "false") .schema(stringSchema) .load("myCsvFilePath.csv") // Convert each non-standard column .withColumn("column1", col("column1").cast(IntegerType)) .withColumn("double_column", regexp_replace(col("double_str"), ",", ".").cast(DoubleType)) .withColumn("date_column", to_date(col("date_str"), "dd.MM.yyyy")) // Optional: Drop the original string columns .drop("double_str", "date_str");
Summary
- Prefer the first approach when possible: using CSV reader options is cleaner and more performant, as parsing happens during the initial read.
- Use the second approach only for complex, inconsistent formats that the built-in options can't handle.
内容的提问来源于stack exchange,提问作者Vladimir Kravets

