You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark正则适配问题:提取指定条件下的苹果数量

问题描述

我在PySpark中有一个名为TEXT的列,数据示例如下:

The sky is red. I have 2 apples and I am fine.
----------------------------------------------
The sky is back. I have 8 apples or I am fine.
----------------------------------------------
The sky is back. she has 8 apples and I am fine.
----------------------------------------------
The hill is red. I have 3 apples and I am fine

需要用F.regexp_extract提取苹果数量,满足以下限制:

  • 仅处理以The sky开头的文本
  • 仅提取I have之后的苹果数量
  • 仅提取后续紧跟and的苹果数量

预期只提取第一条数据中的数字2,但我写的正则(?<=The sky.*?)(?<=I have )(.*?)(?= apples and)在PySpark中报错:SparkRuntimeException: [INVALID_PARAMETER_VALUE] The value of parameter(s) 'regexp' in regexp_extract is invalid,该正则在regexstorm.net测试有效,但PySpark不兼容。

附报错代码示例:

import pyspark.sql.functions as F
from pyspark.sql import DataFrame as SparkDataFrame


regex_pattern = '(?<=The sky.*?)(?<=I have )(.*?)(?= apples and)'

df = spark.createDataFrame(
    [
        ('The sky is red. I have 2 apples and I am fine.', 2),
        ('The sky is back. I have 8 apples or I am fine.', 8),
        ('The sky is back. she has 8 apples and I am fine.', 8), 
        ('The hill is red. I have 3 apples and I am fine', 3),
    ],
    ["TEXT", "APPLES"]
)

df = df.withColumn('Number_Apples', F.regexp_extract(F.col("TEXT"), regex_pattern, 0))

df.display()
解决方案

PySpark基于Java正则引擎,和.NET正则引擎(regexstorm使用的引擎)规则不同,原正则报错的核心原因:

  1. Java正则不支持可变长度的正向预查(lookbehind),比如(?<=The sky.*?)中的.*?是可变长度,不符合要求;
  2. Java正则不支持多个连续的重叠预查。

修改思路是放弃预查,改用捕获组匹配整个符合条件的前缀,直接捕获目标数字:

修改后的正则表达式

^The sky.*?I have (\d+) apples and

正则说明:

  • ^ 匹配文本开头,确保内容以The sky起始;
  • .*? 非贪婪匹配The sky到I have之间的任意内容;
  • (\d+) 捕获数字(苹果数量),作为第一个捕获组;
  • apples and 确保数字后紧跟指定字符串,满足“后续紧跟and”的限制。

修改后的代码

import pyspark.sql.functions as F
from pyspark.sql import DataFrame as SparkDataFrame


# 修改后的正则表达式(注意Python中需要转义反斜杠)
regex_pattern = '^The sky.*?I have (\\d+) apples and'

df = spark.createDataFrame(
    [
        ('The sky is red. I have 2 apples and I am fine.', 2),
        ('The sky is back. I have 8 apples or I am fine.', 8),
        ('The sky is back. she has 8 apples and I am fine.', 8), 
        ('The hill is red. I have 3 apples and I am fine', 3),
    ],
    ["TEXT", "APPLES"]
)

# 提取第一个捕获组(索引为1,索引0是整个正则匹配的内容)
df = df.withColumn('Number_Apples', F.regexp_extract(F.col("TEXT"), regex_pattern, 1))

df.display()

验证结果

  • 第一条数据:成功提取2,符合预期;
  • 第二条数据:因后续是or而非and,匹配失败,返回空字符串;
  • 第三条数据:因是she has而非I have,匹配失败,返回空字符串;
  • 第四条数据:因开头不是The sky,匹配失败,返回空字符串。

内容的提问来源于stack exchange,提问作者maede nasri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 17:57:05