You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark如何识别商品首次出现年份并标记2021年新上市商品

PySpark 实现商品新旧标签计算

实现思路

  • 按商品字段desc分组,计算每个商品的首次出现年份(即分组内year的最小值)
  • 若商品首次出现年份为2021,该商品所有对应行标签设为New,其余所有情况标签设为old

完整代码

from pyspark.sql import functions as F
from pyspark.sql.window import Window

# 构造测试数据
df = spark.createDataFrame(
  [
     ('a','2019')
    ,('a','2020')
    ,('a','2020')
    ,('b','2020')
    ,('b','2019')
    ,('b','2021')
    ,('c','2021')
    ,('a','2021')
    ,('c','2021')
    ,('e','2020')
  ], ['desc', 'year'])

# 定义分组窗口,按商品desc分区
win = Window.partitionBy("desc")
# 计算标签
df_result = df.withColumn("min_year", F.min("year").over(win)) \
              .withColumn("Label", F.when(F.col("min_year") == "2021", "New").otherwise("old")) \
              .drop("min_year") # 移除中间计算列

# 按原数据顺序输出结果(可根据需求省略)
df_result.orderBy(F.monotonically_increasing_id()).show()

输出结果

+----+----+-----+
|desc|year|Label|
+----+----+-----+
|   a|2019|  old|
|   a|2020|  old|
|   a|2020|  old|
|   b|2020|  old|
|   b|2019|  old|
|   b|2021|  old|
|   c|2021|  New|
|   a|2021|  old|
|   c|2021|  New|
|   e|2020|  old|
+----+----+-----+

如果你的数据中year字段为整数类型,只需要将判断条件中的"2021"改为2021即可适配。

内容的提问来源于stack exchange,提问作者naveen kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 00:06:04