Spark Java:无需指定Schema拆分Dataset单列为N列及分隔符处理
Dynamic Splitting to N Columns (No Hardcoded Schema)
Since you want to avoid tying yourself to a fixed schema and handle any number of columns, this is totally doable in Java Spark. Here's a practical implementation:
- First, split your raw combined column into an array using the
|delimiter (note: we escape it with\\since it’s a regex special character). - Grab the number of elements from the split array (using the first row) to know how many columns to generate.
- Dynamically create columns from the array—use your predefined names (like
id,tid, etc.) if they exist, or fall back to generic names if the column count exceeds your list.
Here’s the code:
import org.apache.spark.sql.Dataset; import org.apache.spark.sql.Row; import static org.apache.spark.sql.functions.*; import java.util.ArrayList; import java.util.Arrays; import java.util.List; // Assume your input Dataset is named rawDs with a single column "raw_data" Dataset<Row> splitArrayDs = rawDs.withColumn("split_arr", split(col("raw_data"), "\\|")); // Get total columns from the split array int totalColumns = splitArrayDs.select(size(col("split_arr"))).head().getInt(0); // Predefined column names matching your expected fields List<String> predefinedColNames = Arrays.asList("id", "tid", "mid", "amount", "mname", "desc", "brand", "brandId", "mcc"); // Build list of columns to select List<Column> targetColumns = new ArrayList<>(); for (int i = 0; i < totalColumns; i++) { String colName = (i < predefinedColNames.size()) ? predefinedColNames.get(i) : "col_" + i; targetColumns.add(col("split_arr").getItem(i).alias(colName)); } // Final Dataset with split columns Dataset<Row> finalDs = splitArrayDs.select(targetColumns.toArray(new Column[0]));
Handling desc Columns With Embedded | Using Double Quotes
Absolutely—wrapping the desc field in double quotes is the right way to prevent Spark from splitting on its internal | characters. Here’s how to make it work:
1. If Your Raw Data Already Has Quoted desc Fields
Use a regex that only splits on | characters outside double quotes. The pattern \|(?=(?:[^"]*"[^"]*")*[^"]*$) does exactly this—it matches | only if it’s followed by an even number of double quotes (meaning it’s not inside a quoted field).
Update the split line in the code above to:
Dataset<Row> splitArrayDs = rawDs.withColumn("split_arr", split(col("raw_data"), "\\|(?=(?:[^\"]*\"[^\"]*\")*[^\"]*$)"));
This will keep quoted fields like "DESC|123" as a single element in the split array.
2. If You Need to Add Quotes to the desc Field First
If your raw data doesn’t already quote the desc field, preprocess it to wrap that segment in quotes. For example, if desc is the 6th field (0-based index 5):
// Wrap the desc field in double quotes Dataset<Row> quotedDs = rawDs.withColumn("quoted_raw", regexp_replace( col("raw_data"), "^((\\|[^|]*){5})([^|]*)(\\|.*)$", "$1\"$3\"$4" )); // Then split using the quote-aware regex Dataset<Row> splitArrayDs = quotedDs.withColumn("split_arr", split(col("quoted_raw"), "\\|(?=(?:[^\"]*\"[^\"]*\")*[^\"]*$)"));
Bonus: Use Spark’s CSV Reader (If Applicable)
If your data behaves like a CSV with | as the delimiter, skip manual splitting entirely—Spark’s built-in CSV reader handles quoted fields automatically:
Dataset<Row> finalDs = spark.read() .option("delimiter", "|") .option("quote", "\"") .option("header", false) // Set to true if your data has a header row .csv("path/to/your/data");
This is the most robust approach if your data follows standard CSV formatting rules.
内容的提问来源于stack exchange,提问作者Shibu

