PySpark:如何使用带参数的UDF通过列表达式创建新列
Problem Description
I've defined the following custom UDF to generate a new column in a DataFrame:
import datetime from pyspark.sql.types import DateType def to_date_formatted(date_str, format): if date_str == '' or date_str is None: return None try: dt = datetime.datetime.strptime(date_str, format) except: return None return dt.date() spark.udf.register("to_date_udf", to_date_formatted, DateType())
I can call this function successfully via SQL like:
select to_date_udf(my_date, '%d-%b-%y') as date
which lets me pass a custom format parameter. But when I try to use it with PySpark column expression syntax (like df.with_column("date", to_date_udf('my_date', '%d-%b-%y'))), I get errors. I want to figure out the correct way to implement this.
Note: The built-in
to_datefunction supports optional format parameters in Spark 2.2+, but I'm using Spark 2.0 and can't use that feature. This date conversion example is just for demonstration—my focus is on the general syntax for passing parameters to UDFs, not the specific date conversion implementation.
Solution
The issue here is that you can't directly use the registered UDF's string name in column expressions. Instead, you need to work with the UDF object returned by spark.udf.register, or wrap your Python function using pyspark.sql.functions.udf. Here are two reliable approaches:
Approach 1: Use the UDF object returned during registration
When you call spark.udf.register, it returns a UDF instance that's compatible with PySpark's column API. Just save this return value and use it to call the function with columns and parameters:
import datetime from pyspark.sql.types import DateType from pyspark.sql.functions import col # Define the function and register it, saving the returned UDF object def to_date_formatted(date_str, format): if date_str == '' or date_str is None: return None try: dt = datetime.datetime.strptime(date_str, format) except: return None return dt.date() to_date_udf = spark.udf.register("to_date_udf", to_date_formatted, DateType()) # Now use column expression syntax correctly df = df.withColumn("date", to_date_udf(col("my_date"), "%d-%b-%y"))
Important: Pass columns using col("my_date") or df["my_date"] (not a string literal like 'my_date'). String literals would be treated as raw values instead of references to DataFrame columns.
Approach 2: Wrap the function directly with udf (no SQL registration)
If you don't need to call the UDF via SQL queries, you can skip registering it and directly wrap your function with pyspark.sql.functions.udf:
import datetime from pyspark.sql.functions import udf, col from pyspark.sql.types import DateType def to_date_formatted(date_str, format): if date_str == '' or date_str is None: return None try: dt = datetime.datetime.strptime(date_str, format) except: return None return dt.date() # Create the UDF object without SQL registration to_date_udf = udf(to_date_formatted, DateType()) # Use it in column expressions df = df.withColumn("date", to_date_udf(col("my_date"), "%d-%b-%y"))
Key Takeaways
- Always pass column references (not string literals) when using UDFs with column expressions—this ensures the UDF reads data from the DataFrame columns instead of treating inputs as static values.
- If you register a UDF for SQL use, remember to capture the returned UDF object for use in PySpark's column API; the string name only works in SQL queries.
This syntax works for any UDF that needs custom parameters, not just date conversion use cases.
内容的提问来源于stack exchange,提问作者RobinL

