You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PySpark DataFrame中为连续字母与数字间添加空格

解决PySpark DataFrame中字母与数字连续的分隔问题

我有一个包含文本列的PySpark DataFrame,部分文本中存在字母与数字连续的情况,例如Machine1234需转为Machine 1234、5years需转为5 years。现有DataFrame如下:

+---+--------------------------------------------+
|id |words                                       |
+---+--------------------------------------------+
|0  |This is Spark123 of 5years                  |
|1  |I wish Java DL1234 could use case classes444|
|2  |Data science is  cool321                    |
|3  |Machine345                                  |
+---+--------------------------------------------+

我尝试使用以下代码但未生效:

df2 = temp.select('id',
    F.regexp_replace('words', r'(\d+(\.\d+)?)', ' \1').alias('words'))

期望得到如下输出:

+---+----------------------------------------------+
|id |words                                         |
+---+----------------------------------------------+
|0  |This is Spark 123 of 5 years                  |
|1  |I wish Java DL 1234 could use case classes 444|
|2  |Data science is  cool 321                     |
|3  |Machine 345                                   |
+---+----------------------------------------------+

问题出在哪

原来的正则只会给数字前面加空格,没法处理数字后面跟字母的情况(比如5years),而且没区分字母和数字的边界,得同时覆盖字母后接数字和数字后接字母两种场景才行。

正确的实现方式

有两种可行方案,任选其一即可:

方案1:两次替换分别处理两种场景

import pyspark.sql.functions as F

df2 = temp.select(
    'id',
    F.regexp_replace(
        F.regexp_replace('words', r'([a-zA-Z])(\d)', r'\1 \2'),
        r'(\d)([a-zA-Z])', r'\1 \2'
    ).alias('words')
)

方案2:用零宽度断言一次搞定

import pyspark.sql.functions as F

df2 = temp.select(
    'id',
    F.regexp_replace('words', r'(?<=[a-zA-Z])(?=\d)|(?<=\d)(?=[a-zA-Z])', ' ').alias('words')
)

代码说明

  • 方案1:
    第一个正则匹配字母后跟数字的情况(比如Spark123→Spark 123),在两者之间插入空格;第二个正则匹配数字后跟字母的情况(比如5years→5 years),同样插入空格。
  • 方案2:
    用零宽度断言精准定位字母和数字的边界:(?<=[a-zA-Z])(?=\d) 找字母后面、数字前面的位置;(?<=\d)(?=[a-zA-Z]) 找数字后面、字母前面的位置,然后在这些位置统一插入空格,一次替换完成所有处理。

执行结果

运行上述代码后,得到的DataFrame和期望输出完全一致:

+---+----------------------------------------------+
|id |words                                         |
+---+----------------------------------------------+
|0  |This is Spark 123 of 5 years                  |
|1  |I wish Java DL 1234 could use case classes 444|
|2  |Data science is  cool 321                     |
|3  |Machine 345                                   |
+---+----------------------------------------------+

内容的提问来源于stack exchange,提问作者merkle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 22:15:41