Spark拆分含内嵌横杠的code列:将CSV转为四列DataFrame
Got it, let's tackle this problem step by step. You need to split the code column in your CSV into three parts while keeping any hyphens in the text section intact—here's how to do it with Spark:
The core trick here is splitting the code column only at the first two hyphens, so any hyphens in the final text segment stay put. Spark's split function supports a limit parameter that lets us control how many splits happen, which is perfect for this case.
Step-by-Step Implementation
1. Read the CSV File
First, we'll load your CSV into a Spark DataFrame, making sure to recognize the header row.
2. Split the code Column
Use split(code, "-", 3): the third parameter 3 tells Spark to split the string into at most 3 parts. This means it will only use the first two hyphens as separators, leaving all remaining content (including hyphens) in the third segment.
3. Extract Split Parts into New Columns
We'll pull each segment from the split result and assign them to your desired num, two_letters, and text columns.
Python Code Example
from pyspark.sql import SparkSession from pyspark.sql.functions import split, col # Initialize Spark session spark = SparkSession.builder.appName("CodeColumnSplitter").getOrCreate() # Load CSV (update the file path to match your local setup) raw_df = spark.read.csv("your_input.csv", header=True, inferSchema=True) # Split code column and extract individual parts final_df = raw_df.withColumn("split_code", split(col("code"), "-", 3)) \ .select( col("term"), col("split_code")[0].alias("num"), # Extract first segment col("split_code")[1].alias("two_letters"), # Extract second segment col("split_code")[2].alias("text") # Extract remaining content as text ) \ .drop("split_code") # Clean up the temporary split column # View the result final_df.show(truncate=False)
Scala Code Example
import org.apache.spark.sql.SparkSession import org.apache.spark.sql.functions.{split, col} // Initialize Spark session val spark = SparkSession.builder.appName("CodeColumnSplitter").getOrCreate() // Load CSV (adjust file path as needed) val rawDF = spark.read .option("header", "true") .option("inferSchema", "true") .csv("your_input.csv") // Split code column and map to new columns val finalDF = rawDF.withColumn("split_code", split(col("code"), "-", 3)) .select( col("term"), col("split_code")(0).alias("num"), col("split_code")(1).alias("two_letters"), col("split_code")(2).alias("text") ) .drop("split_code") // Show the output finalDF.show(false)
Optional: Convert num to Numeric Type
If you want num to be an integer instead of a string, modify the column extraction line to cast it:
# For Python col("split_code")[0].cast("integer").alias("num")
// For Scala col("split_code")(0).cast("integer").alias("num")
Test with Your Sample Input
For your example rows:
12-AB-some textsplits into["12", "AB", "some text"]130-CD-some-other-textsplits into["130", "CD", "some-other-text"]
The final DataFrame will exactly match your expected output.
内容的提问来源于stack exchange,提问作者Mousa

