You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark拆分含内嵌横杠的code列:将CSV转为四列DataFrame

Got it, let's tackle this problem step by step. You need to split the code column in your CSV into three parts while keeping any hyphens in the text section intact—here's how to do it with Spark:

Solution Overview

The core trick here is splitting the code column only at the first two hyphens, so any hyphens in the final text segment stay put. Spark's split function supports a limit parameter that lets us control how many splits happen, which is perfect for this case.

Step-by-Step Implementation

1. Read the CSV File

First, we'll load your CSV into a Spark DataFrame, making sure to recognize the header row.

2. Split the code Column

Use split(code, "-", 3): the third parameter 3 tells Spark to split the string into at most 3 parts. This means it will only use the first two hyphens as separators, leaving all remaining content (including hyphens) in the third segment.

3. Extract Split Parts into New Columns

We'll pull each segment from the split result and assign them to your desired num, two_letters, and text columns.


Python Code Example

from pyspark.sql import SparkSession
from pyspark.sql.functions import split, col

# Initialize Spark session
spark = SparkSession.builder.appName("CodeColumnSplitter").getOrCreate()

# Load CSV (update the file path to match your local setup)
raw_df = spark.read.csv("your_input.csv", header=True, inferSchema=True)

# Split code column and extract individual parts
final_df = raw_df.withColumn("split_code", split(col("code"), "-", 3)) \
    .select(
        col("term"),
        col("split_code")[0].alias("num"),  # Extract first segment
        col("split_code")[1].alias("two_letters"),  # Extract second segment
        col("split_code")[2].alias("text")  # Extract remaining content as text
    ) \
    .drop("split_code")  # Clean up the temporary split column

# View the result
final_df.show(truncate=False)

Scala Code Example

import org.apache.spark.sql.SparkSession
import org.apache.spark.sql.functions.{split, col}

// Initialize Spark session
val spark = SparkSession.builder.appName("CodeColumnSplitter").getOrCreate()

// Load CSV (adjust file path as needed)
val rawDF = spark.read
    .option("header", "true")
    .option("inferSchema", "true")
    .csv("your_input.csv")

// Split code column and map to new columns
val finalDF = rawDF.withColumn("split_code", split(col("code"), "-", 3))
    .select(
        col("term"),
        col("split_code")(0).alias("num"),
        col("split_code")(1).alias("two_letters"),
        col("split_code")(2).alias("text")
    )
    .drop("split_code")

// Show the output
finalDF.show(false)

Optional: Convert num to Numeric Type

If you want num to be an integer instead of a string, modify the column extraction line to cast it:

# For Python
col("split_code")[0].cast("integer").alias("num")
// For Scala
col("split_code")(0).cast("integer").alias("num")

Test with Your Sample Input

For your example rows:

  • 12-AB-some text splits into ["12", "AB", "some text"]
  • 130-CD-some-other-text splits into ["130", "CD", "some-other-text"]

The final DataFrame will exactly match your expected output.

内容的提问来源于stack exchange,提问作者Mousa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:31:29