You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在PySpark中创建可直接返回正则提取结果列的自定义函数

Modify the Function to Return Only the Extracted Column

To achieve your desired usage pattern (df['new_column'] = extract_strings(df, 'text')), you need to adjust the function to return the regex-extracted result as a PySpark Column object instead of the entire DataFrame. Here's how:

Modified Function Code

from pyspark.sql import functions as F

def extract_strings(dataframe_selected, column_selected):
    # Directly return the regex extraction result as a Column expression
    return F.regexp_extract(dataframe_selected[column_selected], r"([a-zA-Z]+)", 0)

Key Changes Explained

  • Removed withColumn: Instead of creating a new DataFrame with an added column, we directly return the output of F.regexp_extract(), which is a Column object representing the extracted values.
  • Simplified Return Value: This Column can be directly assigned to a new column in your original DataFrame using the pandas-style syntax you prefer.

Usage Example

# Assign the extracted values to a new column in your DataFrame
df['new_column'] = extract_strings(df, 'text')

This approach keeps your code concise and aligns with the assignment pattern you want to use.

内容的提问来源于stack exchange,提问作者Hernandes Matias Junior

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 14:32:45