基于PySpark从DataFrame字符串列提取关键词前的单词
Got it, let's tackle this problem step by step. You want to extract the word immediately preceding the Key_word in the Text column, right? Here's a clean, efficient approach using PySpark's built-in functions (no slow UDFs required):
Full Solution Code
from pyspark.sql import SparkSession from pyspark.sql.functions import split, array_position, element_at, when, col # Initialize Spark session spark = SparkSession.builder.appName("PrecedingWordExtractor").getOrCreate() # Your sample data data = [ ("First random text tree cheese cat", "tree"), ("Second random text apple pie three", "text"), ("Third random text burger food brain", "brain"), ("Fourth random text nothing thing chips", "random"), ] # Create initial DataFrame df = spark.createDataFrame(data, ["Text", "Key_word"]) # Add the preceding word column result_df = df.withColumn( "word_array", split(col("Text"), "\\s+") # Split text into words (handles multiple spaces) ).withColumn( "key_position", array_position(col("word_array"), col("Key_word")) # Find position of key word (1-indexed) ).withColumn( "Preceding_word", when( col("key_position") > 1, # Only if key isn't the first word element_at(col("word_array"), col("key_position") - 1) ).otherwise(None) # Return null if key is first word ).drop("word_array", "key_position") # Clean up intermediate columns # Show the result result_df.show(truncate=False)
Output
+----------------------------------------+----------+-------------+ |Text |Key_word |Preceding_word| +----------------------------------------+----------+-------------+ |First random text tree cheese cat |tree |text | |Second random text apple pie three |text |random | |Third random text burger food brain |brain |food | |Fourth random text nothing thing chips |random |Fourth | +----------------------------------------+----------+-------------+
How It Works
Let's break down each step:
split(col("Text"), "\\s+"): Splits theTextcolumn into an array of words. Using\\s+instead of a single space ensures we handle cases where there are multiple spaces between words.array_position(col("word_array"), col("Key_word")): Finds the position ofKey_wordin the word array. Important note: PySpark'sarray_positionuses 1-based indexing (unlike Python's 0-based), so the first word is position 1.when(...): Checks if the key word isn't the first word (position > 1). If true, it useselement_atto grab the word at positionkey_position - 1(the one right before the key). If the key is first, it returnsnull.drop(...): Removes the intermediate columns we created to keep the final DataFrame clean.
Edge Case Handling
If you want to replace null with a custom value (like "No preceding word") instead of leaving it blank, just modify the otherwise clause:
.when( col("key_position") > 1, element_at(col("word_array"), col("key_position") - 1) ).otherwise("No preceding word")
内容的提问来源于stack exchange,提问作者Anna
相关产品推荐
相关产品推荐

