You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark pandas Series.replace正则替换抛NotImplementedError问题

问题

使用PySpark 3.3.1版本的PySpark pandas时,尝试编写resolve_abbreviations函数处理pyspark.pandas.series.Series类型数据,用正则表达式将缩写替换为完整表述,但调用时抛出NotImplementedError,提示'replace currently not support for regex'。但查阅官方文档,Series.replace函数标注支持正则替换,请问问题出在哪里?

函数代码

def resolve_abbreviations(job_list: pspd.Series) -> pspd.Series:
    """
    The job titles contain a lot of abbreviations for common terms.
    We write them out to create a more standardized job title list.

    :param job_list: df.SchoneFunctie during processing steps
    :return: SchoneFunctie where abbreviations are written out in words
    """
    abbreviations_dict = {
        "1e": "eerste",
        "1ste": "eerste",
        "2e": "tweede",
        "2de": "tweede",
        "3e": "derde",
        "3de": "derde",
        "ceo": "chief executive officer",
        "cfo": "chief financial officer",
        "coo": "chief operating officer",
        "cto": "chief technology officer",
        "sr": "senior",
        "tech": "technisch",
        "zw": "zelfstandig werkend"
    }

    #Create a list of abbreviations
    abbreviations_pob = list(abbreviations_dict.keys())

    #For each abbreviation in this list
    for abb in abbreviations_pob:
        # define patterns to look for
        patterns = [fr'((?<=( ))|(?<=(^))|(?<=(\\))|(?<=(\())){abb}((?=( ))|(?=(\\))|(?=($))|(?=(\))))',
                    fr'{abb}\.']
        # actual recoding of abbreviations to written out form
        value_to_replace = abbreviations_dict[abb]
        for patt in patterns:
            job_list = job_list.replace(to_replace=fr'{patt}', value=f'{value_to_replace} ', regex=True)

    return job_list

调用代码

df['CleanedUp'] = resolve_abbreviations(df['SchoneFunctie'])

报错信息

Traceback (most recent call last):
  File "C:\Program Files\JetBrains\PyCharm 2021.3\plugins\python\helpers\pydev\pydevd.py", line 1496, in _exec
    pydev_imports.execfile(file, globals, locals)  # execute the script
  File "C:\Program Files\JetBrains\PyCharm 2021.3\plugins\python\helpers\pydev\_pydev_imps\_pydev_execfile.py", line 18, in execfile
    exec(compile(contents+"\n", file, 'exec'), glob, loc)
  File "C:\path_to_python_file\python_file.py", line 180, in <module>
    df['SchoneFunctie'] = resolve_abbreviations(df['SchoneFunctie'])
  File "C:\path_to_python_file\python_file.py", line 164, in resolve_abbreviations
    job_list = job_list.replace(to_replace=fr'{patt}', value=f'{value_to_replace} ', regex=True)
  File "C:\Users\MyUser\.conda\envs\Anaconda3.9\lib\site-packages\pyspark\pandas\series.py", line 4492, in replace
    raise NotImplementedError("replace currently not support for regex")
NotImplementedError: replace currently not support for regex
python-BaseException

问题原因与解决方案

原因

PySpark 3.3.1版本中,pyspark.pandas.Series.replace方法的regex=True参数并未实际实现。查看该版本的series.py源码(对应报错中的第4492行),会发现当传入regex=True时,代码直接抛出NotImplementedError。官方文档可能存在更新滞后,或是混淆了原生pandas与PySpark pandas的API差异。

解决方案

改用PySpark原生的regexp_replace函数处理,该函数完全支持正则替换,且能与PySpark pandas的Series兼容。修改后的函数如下:

from pyspark.sql import functions as F
import pyspark.pandas as pspd

def resolve_abbreviations(job_list: pspd.Series) -> pspd.Series:
    """
    将职位名称中的缩写替换为完整表述,生成标准化的职位列表
    :param job_list: 处理中的df.SchoneFunctie列
    :return: 替换缩写后的SchoneFunctie列
    """
    abbreviations_dict = {
        "1e": "eerste",
        "1ste": "eerste",
        "2e": "tweede",
        "2de": "tweede",
        "3e": "derde",
        "3de": "derde",
        "ceo": "chief executive officer",
        "cfo": "chief financial officer",
        "coo": "chief operating officer",
        "cto": "chief technology officer",
        "sr": "senior",
        "tech": "technisch",
        "zw": "zelfstandig werkend"
    }

    # 将PySpark pandas Series转换为Spark Column
    spark_col = job_list.to_spark()
    
    for abb, full_term in abbreviations_dict.items():
        # 构建两个匹配模式
        pattern_boundary = fr'((?<=( ))|(?<=(^))|(?<=(\\))|(?<=(\())){abb}((?=( ))|(?=(\\))|(?=($))|(?=(\))))'
        pattern_dot = fr'{abb}\.'
        
        # 执行正则替换
        spark_col = F.regexp_replace(spark_col, pattern_boundary, f"{full_term} ")
        spark_col = F.regexp_replace(spark_col, pattern_dot, f"{full_term} ")
    
    # 将处理后的Spark Column转回PySpark pandas Series
    return pspd.Series(spark_col)

说明

  • 该方案通过to_spark()将PySpark pandas Series转换为Spark原生Column,利用F.regexp_replace完成正则替换,最后再转回PySpark pandas Series,完全适配原有代码的调用逻辑。
  • 相比原循环调用replace的方式,Spark原生函数的执行效率更高,更适合大数据场景。

内容的提问来源于stack exchange,提问作者Psychotechnopath

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 07:07:49