如何在PySpark DataFrame中为连续字母与数字间添加空格
解决PySpark DataFrame中字母与数字连续的分隔问题
我有一个包含文本列的PySpark DataFrame,部分文本中存在字母与数字连续的情况,例如Machine1234需转为Machine 1234、5years需转为5 years。现有DataFrame如下:
+---+--------------------------------------------+ |id |words | +---+--------------------------------------------+ |0 |This is Spark123 of 5years | |1 |I wish Java DL1234 could use case classes444| |2 |Data science is cool321 | |3 |Machine345 | +---+--------------------------------------------+
我尝试使用以下代码但未生效:
df2 = temp.select('id', F.regexp_replace('words', r'(\d+(\.\d+)?)', ' \1').alias('words'))
期望得到如下输出:
+---+----------------------------------------------+ |id |words | +---+----------------------------------------------+ |0 |This is Spark 123 of 5 years | |1 |I wish Java DL 1234 could use case classes 444| |2 |Data science is cool 321 | |3 |Machine 345 | +---+----------------------------------------------+
问题出在哪
原来的正则只会给数字前面加空格,没法处理数字后面跟字母的情况(比如5years),而且没区分字母和数字的边界,得同时覆盖字母后接数字和数字后接字母两种场景才行。
正确的实现方式
有两种可行方案,任选其一即可:
方案1:两次替换分别处理两种场景
import pyspark.sql.functions as F df2 = temp.select( 'id', F.regexp_replace( F.regexp_replace('words', r'([a-zA-Z])(\d)', r'\1 \2'), r'(\d)([a-zA-Z])', r'\1 \2' ).alias('words') )
方案2:用零宽度断言一次搞定
import pyspark.sql.functions as F df2 = temp.select( 'id', F.regexp_replace('words', r'(?<=[a-zA-Z])(?=\d)|(?<=\d)(?=[a-zA-Z])', ' ').alias('words') )
代码说明
- 方案1:
第一个正则匹配字母后跟数字的情况(比如Spark123→Spark 123),在两者之间插入空格;第二个正则匹配数字后跟字母的情况(比如5years→5 years),同样插入空格。 - 方案2:
用零宽度断言精准定位字母和数字的边界:(?<=[a-zA-Z])(?=\d)找字母后面、数字前面的位置;(?<=\d)(?=[a-zA-Z])找数字后面、字母前面的位置,然后在这些位置统一插入空格,一次替换完成所有处理。
执行结果
运行上述代码后,得到的DataFrame和期望输出完全一致:
+---+----------------------------------------------+ |id |words | +---+----------------------------------------------+ |0 |This is Spark 123 of 5 years | |1 |I wish Java DL 1234 could use case classes 444| |2 |Data science is cool 321 | |3 |Machine 345 | +---+----------------------------------------------+
内容的提问来源于stack exchange,提问作者merkle
相关产品推荐
相关产品推荐

