PySpark正则适配问题:提取指定条件下的苹果数量
问题描述
我在PySpark中有一个名为TEXT的列,数据示例如下:
The sky is red. I have 2 apples and I am fine. ---------------------------------------------- The sky is back. I have 8 apples or I am fine. ---------------------------------------------- The sky is back. she has 8 apples and I am fine. ---------------------------------------------- The hill is red. I have 3 apples and I am fine
需要用F.regexp_extract提取苹果数量,满足以下限制:
- 仅处理以
The sky开头的文本 - 仅提取
I have之后的苹果数量 - 仅提取后续紧跟
and的苹果数量
预期只提取第一条数据中的数字2,但我写的正则(?<=The sky.*?)(?<=I have )(.*?)(?= apples and)在PySpark中报错:SparkRuntimeException: [INVALID_PARAMETER_VALUE] The value of parameter(s) 'regexp' in regexp_extract is invalid,该正则在regexstorm.net测试有效,但PySpark不兼容。
附报错代码示例:
import pyspark.sql.functions as F from pyspark.sql import DataFrame as SparkDataFrame regex_pattern = '(?<=The sky.*?)(?<=I have )(.*?)(?= apples and)' df = spark.createDataFrame( [ ('The sky is red. I have 2 apples and I am fine.', 2), ('The sky is back. I have 8 apples or I am fine.', 8), ('The sky is back. she has 8 apples and I am fine.', 8), ('The hill is red. I have 3 apples and I am fine', 3), ], ["TEXT", "APPLES"] ) df = df.withColumn('Number_Apples', F.regexp_extract(F.col("TEXT"), regex_pattern, 0)) df.display()
解决方案
PySpark基于Java正则引擎,和.NET正则引擎(regexstorm使用的引擎)规则不同,原正则报错的核心原因:
- Java正则不支持可变长度的正向预查(lookbehind),比如
(?<=The sky.*?)中的.*?是可变长度,不符合要求; - Java正则不支持多个连续的重叠预查。
修改思路是放弃预查,改用捕获组匹配整个符合条件的前缀,直接捕获目标数字:
修改后的正则表达式
^The sky.*?I have (\d+) apples and
正则说明:
^匹配文本开头,确保内容以The sky起始;.*?非贪婪匹配The sky到I have之间的任意内容;(\d+)捕获数字(苹果数量),作为第一个捕获组;apples and确保数字后紧跟指定字符串,满足“后续紧跟and”的限制。
修改后的代码
import pyspark.sql.functions as F from pyspark.sql import DataFrame as SparkDataFrame # 修改后的正则表达式(注意Python中需要转义反斜杠) regex_pattern = '^The sky.*?I have (\\d+) apples and' df = spark.createDataFrame( [ ('The sky is red. I have 2 apples and I am fine.', 2), ('The sky is back. I have 8 apples or I am fine.', 8), ('The sky is back. she has 8 apples and I am fine.', 8), ('The hill is red. I have 3 apples and I am fine', 3), ], ["TEXT", "APPLES"] ) # 提取第一个捕获组(索引为1,索引0是整个正则匹配的内容) df = df.withColumn('Number_Apples', F.regexp_extract(F.col("TEXT"), regex_pattern, 1)) df.display()
验证结果
- 第一条数据:成功提取
2,符合预期; - 第二条数据:因后续是
or而非and,匹配失败,返回空字符串; - 第三条数据:因是
she has而非I have,匹配失败,返回空字符串; - 第四条数据:因开头不是
The sky,匹配失败,返回空字符串。
内容的提问来源于stack exchange,提问作者maede nasri
相关产品推荐
相关产品推荐

