You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark:如何使用带参数的UDF通过列表达式创建新列

How to Pass Custom Parameters to a Registered UDF Using PySpark Column Expression Syntax?

Problem Description

I've defined the following custom UDF to generate a new column in a DataFrame:

import datetime
from pyspark.sql.types import DateType

def to_date_formatted(date_str, format):
    if date_str == '' or date_str is None:
        return None
    try:
        dt = datetime.datetime.strptime(date_str, format)
    except:
        return None
    return dt.date()

spark.udf.register("to_date_udf", to_date_formatted, DateType())

I can call this function successfully via SQL like:

select to_date_udf(my_date, '%d-%b-%y') as date

which lets me pass a custom format parameter. But when I try to use it with PySpark column expression syntax (like df.with_column("date", to_date_udf('my_date', '%d-%b-%y'))), I get errors. I want to figure out the correct way to implement this.

Note: The built-in to_date function supports optional format parameters in Spark 2.2+, but I'm using Spark 2.0 and can't use that feature. This date conversion example is just for demonstration—my focus is on the general syntax for passing parameters to UDFs, not the specific date conversion implementation.

Solution

The issue here is that you can't directly use the registered UDF's string name in column expressions. Instead, you need to work with the UDF object returned by spark.udf.register, or wrap your Python function using pyspark.sql.functions.udf. Here are two reliable approaches:

Approach 1: Use the UDF object returned during registration

When you call spark.udf.register, it returns a UDF instance that's compatible with PySpark's column API. Just save this return value and use it to call the function with columns and parameters:

import datetime
from pyspark.sql.types import DateType
from pyspark.sql.functions import col

# Define the function and register it, saving the returned UDF object
def to_date_formatted(date_str, format):
    if date_str == '' or date_str is None:
        return None
    try:
        dt = datetime.datetime.strptime(date_str, format)
    except:
        return None
    return dt.date()

to_date_udf = spark.udf.register("to_date_udf", to_date_formatted, DateType())

# Now use column expression syntax correctly
df = df.withColumn("date", to_date_udf(col("my_date"), "%d-%b-%y"))

Important: Pass columns using col("my_date") or df["my_date"] (not a string literal like 'my_date'). String literals would be treated as raw values instead of references to DataFrame columns.

Approach 2: Wrap the function directly with udf (no SQL registration)

If you don't need to call the UDF via SQL queries, you can skip registering it and directly wrap your function with pyspark.sql.functions.udf:

import datetime
from pyspark.sql.functions import udf, col
from pyspark.sql.types import DateType

def to_date_formatted(date_str, format):
    if date_str == '' or date_str is None:
        return None
    try:
        dt = datetime.datetime.strptime(date_str, format)
    except:
        return None
    return dt.date()

# Create the UDF object without SQL registration
to_date_udf = udf(to_date_formatted, DateType())

# Use it in column expressions
df = df.withColumn("date", to_date_udf(col("my_date"), "%d-%b-%y"))

Key Takeaways

  • Always pass column references (not string literals) when using UDFs with column expressions—this ensures the UDF reads data from the DataFrame columns instead of treating inputs as static values.
  • If you register a UDF for SQL use, remember to capture the returned UDF object for use in PySpark's column API; the string name only works in SQL queries.

This syntax works for any UDF that needs custom parameters, not just date conversion use cases.

内容的提问来源于stack exchange,提问作者RobinL

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:10:46