You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Snowpark Snowflake自定义UDF输入及返回类型配置报错求助

Snowpark自定义UDF类型配置报错解决方案

问题根源

你用@udf(标量UDF装饰器)时直接指定PandasDataFrame作为输入/返回类型,这是错误的:

  • 标量UDF仅支持处理单行的单个或多个字段,不能直接接收/返回整个DataFrame
  • Snowpark不支持PandasDataFrame[int, float, float]这种泛型写法定义数据类型,需要用官方的StructType声明Schema

正确解决方案

根据需求选择对应的UDF类型:

1. 处理批量行的向量UDF(Pandas UDF)

如果需要处理整表的Pandas DataFrame,使用@pandas_udf装饰器,同时用Snowpark的Schema类型定义输入输出结构:

import pandas as pd
from snowflake.snowpark.functions import pandas_udf
from snowflake.snowpark.types import StructType, StructField, IntegerType, FloatType, StringType

# 定义输入列的Schema
input_schema = StructType([
    StructField("COL1", IntegerType()),
    StructField("COL2", FloatType()),
    StructField("COL3", FloatType())
])

# 定义输出列的Schema
output_schema = StructType([
    StructField("COL1", IntegerType()),
    StructField("RESULT_COL", StringType())
])

@pandas_udf(output_schema, input_types=[input_schema])
def new(df: pd.DataFrame) -> pd.DataFrame:
    # 写入你的数据处理逻辑
    df["RESULT_COL"] = df["COL2"].astype(str) + "_" + df["COL3"].astype(str)
    # 返回指定列的DataFrame
    return df[["COL1", "RESULT_COL"]]

2. 处理单行数据的标量UDF

如果只是处理单行的多个字段,用@udf装饰器,直接指定单个字段的类型:

from snowflake.snowpark.functions import udf
from snowflake.snowpark.types import IntegerType, FloatType, StringType

@udf(input_types=[IntegerType(), FloatType(), FloatType], return_type=StringType())
def single_row_process(col1: int, col2: float, col3: float) -> str:
    # 单行数据处理逻辑
    return f"{col1}_processed_{col2 + col3}"

内容的提问来源于stack exchange,提问作者Abhishek Patil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 15:45:35