如何在PySpark中基于索引奇偶性创建label列(poi1/poi2)
PySpark实现按行索引奇偶性生成label列
PySpark DataFrame没有Pandas那样的原生行索引,所以得先生成连续的行序号,再根据序号的奇偶性生成目标label列。下面提供两种常用实现方法:
方法1:用monotonically_increasing_id()生成全局递增ID
适合不需要严格连续行号的场景,ID全局唯一且递增,分布式环境下也能正常使用:
from pyspark.sql.functions import when, monotonically_increasing_id # 给DataFrame添加自增ID列 df = df.withColumn("row_idx", monotonically_increasing_id()) # 根据row_idx+1的奇偶性赋值label df = df.withColumn( "label", when((df.row_idx + 1) % 2 != 0, "poi1").otherwise("poi2") ) # 不需要临时列的话可以删掉 df = df.drop("row_idx")
方法2:用row_number()生成严格连续行号
如果需要严格按数据顺序生成连续行号,就用窗口函数:
from pyspark.sql.functions import when, row_number from pyspark.sql.window import Window # 定义无分区窗口,如需按特定列排序,在orderBy里指定,比如orderBy("date_id", "minute") window_spec = Window.orderBy([]) # 添加连续行号列(这里的row_idx就是原需求里的索引+1) df = df.withColumn("row_idx", row_number().over(window_spec)) # 根据行号奇偶性生成label df = df.withColumn( "label", when(df.row_idx % 2 != 0, "poi1").otherwise("poi2") ) # 删除临时行号列 df = df.drop("row_idx")
测试结果
用你提供的输入数据验证,两种方法都会输出符合预期的结果:
index date_id year month day hour minute label 0 156454 20200801 2021 12 31 12 38 poi1 1 156454 20200801 2021 12 31 12 39 poi2
内容的提问来源于stack exchange,提问作者Nabih Bawazir
相关产品推荐
相关产品推荐

