You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PySpark中从现有列右侧截取可变长度字符生成新列?

PySpark提取Product列最后一个连字符后的数字生成UPC列

问题场景

现有PySpark DataFrame结构如下:

ProductName
abcd - 12abcd
xyz - 123543xyz

需要生成新列UPC,仅保留Product列中最后一个连字符右侧的数字(可按需去除前后空格)。

解决方案

方法1:split + element_at(Spark 2.4+推荐)

用split按连字符分割字符串,element_at直接取最后一个分割元素,配合trim清理空格:

from pyspark.sql.functions import split, element_at, trim, col

df = df.withColumn("UPC", trim(element_at(split(col("Product"), "-"), -1)))

方法2:正则表达式提取(适配复杂格式)

用regexp_extract精准匹配最后一个连字符后的数字部分,灵活处理格式差异:

from pyspark.sql.functions import regexp_extract, trim, col

# 提取最后一个'-'后的数字,自动忽略前置空格
df = df.withColumn("UPC", trim(regexp_extract(col("Product"), r'-(\s*\d+)$', 1)))

方法3:修复你原有的substring方案

你之前报错TypeError: Column is not iterable,是因为在substring中直接用df4.LastHyphen引用列,而非使用Column对象。修正后代码:

from pyspark.sql.functions import col, length, locate, reverse, substring, trim

df4 = df3.withColumn("LastHyphen", length(col("Product")) - locate('-', reverse(col("Product"))))
# 用col()引用列对象,且substring起始位置需+1(PySpark中substring从1开始计数)
df4 = df4.withColumn("UPC", trim(substring(col("Product"), col("LastHyphen") + 1, length(col("Product")) - col("LastHyphen"))))

最终输出

处理后DataFrame如下:

ProductUPC
abcd - 1212
xyz - 123543123543

内容的提问来源于stack exchange,提问作者Ben Perner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 21:55:08