在PySpark中创建可直接返回正则提取结果列的自定义函数
Modify the Function to Return Only the Extracted Column
To achieve your desired usage pattern (df['new_column'] = extract_strings(df, 'text')), you need to adjust the function to return the regex-extracted result as a PySpark Column object instead of the entire DataFrame. Here's how:
Modified Function Code
from pyspark.sql import functions as F def extract_strings(dataframe_selected, column_selected): # Directly return the regex extraction result as a Column expression return F.regexp_extract(dataframe_selected[column_selected], r"([a-zA-Z]+)", 0)
Key Changes Explained
- Removed
withColumn: Instead of creating a new DataFrame with an added column, we directly return the output ofF.regexp_extract(), which is a Column object representing the extracted values. - Simplified Return Value: This Column can be directly assigned to a new column in your original DataFrame using the pandas-style syntax you prefer.
Usage Example
# Assign the extracted values to a new column in your DataFrame df['new_column'] = extract_strings(df, 'text')
This approach keeps your code concise and aligns with the assignment pattern you want to use.
内容的提问来源于stack exchange,提问作者Hernandes Matias Junior
相关产品推荐
相关产品推荐

