You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于PySpark从DataFrame字符串列提取关键词前的单词

Got it, let's tackle this problem step by step. You want to extract the word immediately preceding the Key_word in the Text column, right? Here's a clean, efficient approach using PySpark's built-in functions (no slow UDFs required):

Full Solution Code

from pyspark.sql import SparkSession
from pyspark.sql.functions import split, array_position, element_at, when, col

# Initialize Spark session
spark = SparkSession.builder.appName("PrecedingWordExtractor").getOrCreate()

# Your sample data
data = [
    ("First random text tree cheese cat", "tree"),
    ("Second random text apple pie three", "text"),
    ("Third random text burger food brain", "brain"),
    ("Fourth random text nothing thing chips", "random"),
]

# Create initial DataFrame
df = spark.createDataFrame(data, ["Text", "Key_word"])

# Add the preceding word column
result_df = df.withColumn(
    "word_array", split(col("Text"), "\\s+")  # Split text into words (handles multiple spaces)
).withColumn(
    "key_position", array_position(col("word_array"), col("Key_word"))  # Find position of key word (1-indexed)
).withColumn(
    "Preceding_word",
    when(
        col("key_position") > 1,  # Only if key isn't the first word
        element_at(col("word_array"), col("key_position") - 1)
    ).otherwise(None)  # Return null if key is first word
).drop("word_array", "key_position")  # Clean up intermediate columns

# Show the result
result_df.show(truncate=False)

Output

+----------------------------------------+----------+-------------+
|Text                                    |Key_word  |Preceding_word|
+----------------------------------------+----------+-------------+
|First random text tree cheese cat       |tree      |text         |
|Second random text apple pie three      |text      |random       |
|Third random text burger food brain     |brain     |food         |
|Fourth random text nothing thing chips  |random    |Fourth       |
+----------------------------------------+----------+-------------+

How It Works

Let's break down each step:

  • split(col("Text"), "\\s+"): Splits the Text column into an array of words. Using \\s+ instead of a single space ensures we handle cases where there are multiple spaces between words.
  • array_position(col("word_array"), col("Key_word")): Finds the position of Key_word in the word array. Important note: PySpark's array_position uses 1-based indexing (unlike Python's 0-based), so the first word is position 1.
  • when(...): Checks if the key word isn't the first word (position > 1). If true, it uses element_at to grab the word at position key_position - 1 (the one right before the key). If the key is first, it returns null.
  • drop(...): Removes the intermediate columns we created to keep the final DataFrame clean.

Edge Case Handling

If you want to replace null with a custom value (like "No preceding word") instead of leaving it blank, just modify the otherwise clause:

.when(
    col("key_position") > 1,
    element_at(col("word_array"), col("key_position") - 1)
).otherwise("No preceding word")

内容的提问来源于stack exchange,提问作者Anna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:35:53