You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark从字符串各单词中提取指定元音字母的实现语法

Hey there! Let's walk through exactly how to extract vowels (a, e, i, o, u) from each word in a string using Spark. I'll use Scala syntax since your example uses val variables, but the logic translates easily to PySpark too.

Step 1: Set up SparkSession and sample data

First, we need to initialize our SparkSession and create a DataFrame with your sample string:

import org.apache.spark.sql.SparkSession
import org.apache.spark.sql.functions._

// Initialize SparkSession (remove .master("local[*]") in production environments)
val spark = SparkSession.builder()
  .appName("VowelExtractionFromWords")
  .master("local[*]")
  .getOrCreate()

// Create DataFrame with your sample input string
val inputDF = spark.createDataFrame(Seq(
  ("I want to extract letters form each word of the string",)
)).toDF("original_string")

Step 2: Split the string into individual words

We'll use split to break the string into an array of words, then explode to turn each word into its own row—this makes it far easier to process each word individually:

val wordsDF = inputDF.withColumn("word", explode(split(col("original_string"), "\\s+")))

Step 3: Extract vowels from each word

For Spark 3.0 and above, the regexp_extract_all function is perfect here—it returns an array of all matching characters. We'll convert words to lowercase first to handle uppercase vowels (like the "I" in your sample):

// Extract all vowels as an array
val vowelsDF = wordsDF.withColumn("extracted_vowels",
  regexp_extract_all(lower(col("word")), "[aeiou]", 0)
)

// Optional: Convert the array to a concatenated string if you need a flat output
val finalDF = vowelsDF.withColumn("vowels_str", concat_ws("", col("extracted_vowels")))

Step 4: View the results

Run show() to see the output in a readable format:

finalDF.select("word", "extracted_vowels", "vowels_str").show(false)

This will give you output like:

+--------+----------------+----------+
|word    |extracted_vowels|vowels_str|
+--------+----------------+----------+
|I       |[i]             |i         |
|want    |[a]             |a         |
|to      |[o]             |o         |
|extract |[e,a]           |ea        |
|letters |[e,e]           |ee        |
|form    |[o]             |o         |
|each    |[e,a]           |ea        |
|word    |[o]             |o         |
|of      |[o]             |o         |
|the     |[e]             |e         |
|string  |[i]             |i         |
+--------+----------------+----------+

For older Spark versions (pre-3.0)

If you're using a Spark version before 3.0 that doesn't support regexp_extract_all, use a custom UDF instead:

// Define a UDF to filter vowels from a given word
val extractVowelsUDF = udf((word: String) => {
  word.toLowerCase.filter(c => Set('a','e','i','o','u').contains(c)).toArray
})

// Apply the UDF to extract vowels from each word
val vowelsDFLegacy = wordsDF.withColumn("extracted_vowels", extractVowelsUDF(col("word")))

Optional: Aggregate back to the original string

If you want to group results back by the original input string (instead of individual word rows), use groupBy and collect_list:

val aggregatedDF = finalDF.groupBy("original_string")
  .agg(collect_list(struct("word", "vowels_str")).alias("word_vowel_pairs"))

aggregatedDF.show(false)

That's all! This approach is efficient and scalable for large datasets, which is exactly what Spark is built for.

内容的提问来源于stack exchange,提问作者SanTuhi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:38:17