低版本PySpark中pyspark.sql.functions.filter函数的替代实现方案
Solution for Older PySpark Versions (No
filter Function) Got it, since your PySpark version doesn't have the array filter function (it was added in PySpark 3.0), we can use simpler, backward-compatible methods to add both the Status and Code columns you need. Here's a clean implementation that works across older versions:
from pyspark.sql import functions as F # Assuming your original DataFrame `df` is already created with your data/schema out_df = df.withColumn( "Code", # Split the string by dots, grab the second element (0-indexed array, so index 1) F.split("LongColumn", "\\.")[1] ).withColumn( "Status", # Check directly if the original string includes "approved" — no array conversions needed! F.when(F.col("LongColumn").contains("approved"), F.lit("approved")) .otherwise(F.lit("unapproved")) ) # Optional: Remove the original LongColumn if you don't need it anymore out_df = out_df.drop("LongColumn")
Breakdown of the Solution:
- Code Column: The
splitfunction turns your dot-separated string into an array. Since you confirmed the second position (after splitting) always has 3 letters, we just access index 1 of that array directly. - Status Column: Instead of converting the string to an array and filtering (which requires the newer
filterfunction), we use the basiccontains()method on the original string. This is not only compatible with older PySpark versions but also more efficient, as we avoid unnecessary array operations.
Why Your Original Code Threw an Error:
The F.filter function for arrays was introduced in PySpark 3.0, so any version before that won't recognize it. Our approach skips this entirely while delivering the exact result you want.
Matching Your Expected Output:
When you run this code with your sample data, you'll get exactly the table you're looking for:
| Code | Status |
|---|---|
| ABC | approved |
| DEF | unapproved |
| ABC | unapproved |
| GHI | approved |
内容的提问来源于stack exchange,提问作者johnnydoe
相关产品推荐
相关产品推荐

