You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PySpark中基于索引奇偶性创建label列(poi1/poi2)

PySpark实现按行索引奇偶性生成label列

PySpark DataFrame没有Pandas那样的原生行索引,所以得先生成连续的行序号,再根据序号的奇偶性生成目标label列。下面提供两种常用实现方法:

方法1:用monotonically_increasing_id()生成全局递增ID

适合不需要严格连续行号的场景,ID全局唯一且递增,分布式环境下也能正常使用:

from pyspark.sql.functions import when, monotonically_increasing_id

# 给DataFrame添加自增ID列
df = df.withColumn("row_idx", monotonically_increasing_id())

# 根据row_idx+1的奇偶性赋值label
df = df.withColumn(
    "label",
    when((df.row_idx + 1) % 2 != 0, "poi1").otherwise("poi2")
)

# 不需要临时列的话可以删掉
df = df.drop("row_idx")

方法2:用row_number()生成严格连续行号

如果需要严格按数据顺序生成连续行号,就用窗口函数:

from pyspark.sql.functions import when, row_number
from pyspark.sql.window import Window

# 定义无分区窗口,如需按特定列排序,在orderBy里指定,比如orderBy("date_id", "minute")
window_spec = Window.orderBy([])

# 添加连续行号列(这里的row_idx就是原需求里的索引+1)
df = df.withColumn("row_idx", row_number().over(window_spec))

# 根据行号奇偶性生成label
df = df.withColumn(
    "label",
    when(df.row_idx % 2 != 0, "poi1").otherwise("poi2")
)

# 删除临时行号列
df = df.drop("row_idx")

测试结果

用你提供的输入数据验证,两种方法都会输出符合预期的结果:

index   date_id     year    month   day hour    minute  label
0   156454  20200801    2021    12       31    12       38  poi1
1   156454  20200801    2021    12       31    12       39  poi2

内容的提问来源于stack exchange,提问作者Nabih Bawazir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 16:10:32