PySpark pandas Series.replace正则替换抛NotImplementedError问题
问题
使用PySpark 3.3.1版本的PySpark pandas时,尝试编写resolve_abbreviations函数处理pyspark.pandas.series.Series类型数据,用正则表达式将缩写替换为完整表述,但调用时抛出NotImplementedError,提示'replace currently not support for regex'。但查阅官方文档,Series.replace函数标注支持正则替换,请问问题出在哪里?
函数代码
def resolve_abbreviations(job_list: pspd.Series) -> pspd.Series: """ The job titles contain a lot of abbreviations for common terms. We write them out to create a more standardized job title list. :param job_list: df.SchoneFunctie during processing steps :return: SchoneFunctie where abbreviations are written out in words """ abbreviations_dict = { "1e": "eerste", "1ste": "eerste", "2e": "tweede", "2de": "tweede", "3e": "derde", "3de": "derde", "ceo": "chief executive officer", "cfo": "chief financial officer", "coo": "chief operating officer", "cto": "chief technology officer", "sr": "senior", "tech": "technisch", "zw": "zelfstandig werkend" } #Create a list of abbreviations abbreviations_pob = list(abbreviations_dict.keys()) #For each abbreviation in this list for abb in abbreviations_pob: # define patterns to look for patterns = [fr'((?<=( ))|(?<=(^))|(?<=(\\))|(?<=(\())){abb}((?=( ))|(?=(\\))|(?=($))|(?=(\))))', fr'{abb}\.'] # actual recoding of abbreviations to written out form value_to_replace = abbreviations_dict[abb] for patt in patterns: job_list = job_list.replace(to_replace=fr'{patt}', value=f'{value_to_replace} ', regex=True) return job_list
调用代码
df['CleanedUp'] = resolve_abbreviations(df['SchoneFunctie'])
报错信息
Traceback (most recent call last): File "C:\Program Files\JetBrains\PyCharm 2021.3\plugins\python\helpers\pydev\pydevd.py", line 1496, in _exec pydev_imports.execfile(file, globals, locals) # execute the script File "C:\Program Files\JetBrains\PyCharm 2021.3\plugins\python\helpers\pydev\_pydev_imps\_pydev_execfile.py", line 18, in execfile exec(compile(contents+"\n", file, 'exec'), glob, loc) File "C:\path_to_python_file\python_file.py", line 180, in <module> df['SchoneFunctie'] = resolve_abbreviations(df['SchoneFunctie']) File "C:\path_to_python_file\python_file.py", line 164, in resolve_abbreviations job_list = job_list.replace(to_replace=fr'{patt}', value=f'{value_to_replace} ', regex=True) File "C:\Users\MyUser\.conda\envs\Anaconda3.9\lib\site-packages\pyspark\pandas\series.py", line 4492, in replace raise NotImplementedError("replace currently not support for regex") NotImplementedError: replace currently not support for regex python-BaseException
问题原因与解决方案
原因
PySpark 3.3.1版本中,pyspark.pandas.Series.replace方法的regex=True参数并未实际实现。查看该版本的series.py源码(对应报错中的第4492行),会发现当传入regex=True时,代码直接抛出NotImplementedError。官方文档可能存在更新滞后,或是混淆了原生pandas与PySpark pandas的API差异。
解决方案
改用PySpark原生的regexp_replace函数处理,该函数完全支持正则替换,且能与PySpark pandas的Series兼容。修改后的函数如下:
from pyspark.sql import functions as F import pyspark.pandas as pspd def resolve_abbreviations(job_list: pspd.Series) -> pspd.Series: """ 将职位名称中的缩写替换为完整表述,生成标准化的职位列表 :param job_list: 处理中的df.SchoneFunctie列 :return: 替换缩写后的SchoneFunctie列 """ abbreviations_dict = { "1e": "eerste", "1ste": "eerste", "2e": "tweede", "2de": "tweede", "3e": "derde", "3de": "derde", "ceo": "chief executive officer", "cfo": "chief financial officer", "coo": "chief operating officer", "cto": "chief technology officer", "sr": "senior", "tech": "technisch", "zw": "zelfstandig werkend" } # 将PySpark pandas Series转换为Spark Column spark_col = job_list.to_spark() for abb, full_term in abbreviations_dict.items(): # 构建两个匹配模式 pattern_boundary = fr'((?<=( ))|(?<=(^))|(?<=(\\))|(?<=(\())){abb}((?=( ))|(?=(\\))|(?=($))|(?=(\))))' pattern_dot = fr'{abb}\.' # 执行正则替换 spark_col = F.regexp_replace(spark_col, pattern_boundary, f"{full_term} ") spark_col = F.regexp_replace(spark_col, pattern_dot, f"{full_term} ") # 将处理后的Spark Column转回PySpark pandas Series return pspd.Series(spark_col)
说明
- 该方案通过
to_spark()将PySpark pandas Series转换为Spark原生Column,利用F.regexp_replace完成正则替换,最后再转回PySpark pandas Series,完全适配原有代码的调用逻辑。 - 相比原循环调用
replace的方式,Spark原生函数的执行效率更高,更适合大数据场景。
内容的提问来源于stack exchange,提问作者Psychotechnopath
相关产品推荐
相关产品推荐

