Spark从字符串各单词中提取指定元音字母的实现语法
Hey there! Let's walk through exactly how to extract vowels (a, e, i, o, u) from each word in a string using Spark. I'll use Scala syntax since your example uses val variables, but the logic translates easily to PySpark too.
Step 1: Set up SparkSession and sample data
First, we need to initialize our SparkSession and create a DataFrame with your sample string:
import org.apache.spark.sql.SparkSession import org.apache.spark.sql.functions._ // Initialize SparkSession (remove .master("local[*]") in production environments) val spark = SparkSession.builder() .appName("VowelExtractionFromWords") .master("local[*]") .getOrCreate() // Create DataFrame with your sample input string val inputDF = spark.createDataFrame(Seq( ("I want to extract letters form each word of the string",) )).toDF("original_string")
Step 2: Split the string into individual words
We'll use split to break the string into an array of words, then explode to turn each word into its own row—this makes it far easier to process each word individually:
val wordsDF = inputDF.withColumn("word", explode(split(col("original_string"), "\\s+")))
Step 3: Extract vowels from each word
For Spark 3.0 and above, the regexp_extract_all function is perfect here—it returns an array of all matching characters. We'll convert words to lowercase first to handle uppercase vowels (like the "I" in your sample):
// Extract all vowels as an array val vowelsDF = wordsDF.withColumn("extracted_vowels", regexp_extract_all(lower(col("word")), "[aeiou]", 0) ) // Optional: Convert the array to a concatenated string if you need a flat output val finalDF = vowelsDF.withColumn("vowels_str", concat_ws("", col("extracted_vowels")))
Step 4: View the results
Run show() to see the output in a readable format:
finalDF.select("word", "extracted_vowels", "vowels_str").show(false)
This will give you output like:
+--------+----------------+----------+ |word |extracted_vowels|vowels_str| +--------+----------------+----------+ |I |[i] |i | |want |[a] |a | |to |[o] |o | |extract |[e,a] |ea | |letters |[e,e] |ee | |form |[o] |o | |each |[e,a] |ea | |word |[o] |o | |of |[o] |o | |the |[e] |e | |string |[i] |i | +--------+----------------+----------+
For older Spark versions (pre-3.0)
If you're using a Spark version before 3.0 that doesn't support regexp_extract_all, use a custom UDF instead:
// Define a UDF to filter vowels from a given word val extractVowelsUDF = udf((word: String) => { word.toLowerCase.filter(c => Set('a','e','i','o','u').contains(c)).toArray }) // Apply the UDF to extract vowels from each word val vowelsDFLegacy = wordsDF.withColumn("extracted_vowels", extractVowelsUDF(col("word")))
Optional: Aggregate back to the original string
If you want to group results back by the original input string (instead of individual word rows), use groupBy and collect_list:
val aggregatedDF = finalDF.groupBy("original_string") .agg(collect_list(struct("word", "vowels_str")).alias("word_vowel_pairs")) aggregatedDF.show(false)
That's all! This approach is efficient and scalable for large datasets, which is exactly what Spark is built for.
内容的提问来源于stack exchange,提问作者SanTuhi

