如何使用PySpark或Spark SQL实现列间模式匹配并更新指定列值?
Hey there! Let's work through this requirement together. Since all your columns are string types, we'll make sure the logic aligns perfectly with that schema. Below are solutions using both PySpark and Spark SQL:
PySpark Approach
We'll use PySpark's when function to conditionally set the "Good Stuff" column value. We've covered both wildcard and regex pattern matching options to fit different needs:
Wildcard Matching (using like)
If your "pattern match" refers to standard SQL wildcards (e.g., % for any character sequence, _ for a single character), use this code:
from pyspark.sql.functions import when, col # Assume your original DataFrame is named `source_df` updated_df = source_df.withColumn( "Good Stuff", when(col("Column A").like(col("Column C")), "1") # Return string "1" on match .otherwise("(-)") # Return string "(-)" when no match ) # Preview the result updated_df.show()
Regex Matching (using rlike)
If you need full regular expression pattern matching instead of basic wildcards, swap like with rlike:
from pyspark.sql.functions import when, col updated_df = source_df.withColumn( "Good Stuff", when(col("Column A").rlike(col("Column C")), "1") .otherwise("(-)") )
Spark SQL Approach
First, register your DataFrame as a temporary view, then use a CASE WHEN statement to handle the conditional logic:
Step 1: Register Temporary View
-- Replace `source_df` with your actual DataFrame name CREATE OR REPLACE TEMP VIEW your_data_view AS SELECT * FROM source_df;
Step 2: Run the Update Query
Wildcard Matching (using LIKE)
SELECT *, CASE WHEN `Column A` LIKE `Column C` THEN '1' ELSE '(-)' END AS `Good Stuff` FROM your_data_view;
Regex Matching (using RLIKE)
SELECT *, CASE WHEN `Column A` RLIKE `Column C` THEN '1' ELSE '(-)' END AS `Good Stuff` FROM your_data_view;
A quick reminder: Since all columns are string types, we're returning string values ('1' and '(-)') to keep consistency with your schema.
内容的提问来源于stack exchange,提问作者anuj

