如何在PySpark中实现与pandas的pd.offsets.MonthEnd(x)等效的功能
PySpark实现pandas pd.offsets.MonthEnd(x)功能的方案
要完全匹配pd.offsets.MonthEnd(x)的效果,需要兼容普通日期和月末日期两种边界场景,全部使用PySpark内置函数实现,性能远高于自定义UDF,适合大数据量场景:
- 若给定日期本身是当月月末:偏移x个月后取新日期的当月月末
- 若给定日期不是当月月末:偏移x-1个月后取新日期的当月月末
完整实现代码
from pyspark.sql import SparkSession from pyspark.sql.functions import last_day, add_months, when, col # 初始化SparkSession(已初始化可跳过) spark = SparkSession.builder.appName("month_end_offset").getOrCreate() # 示例测试数据,包含普通日期、月末日期两种场景 df = spark.createDataFrame( [("2010-01-15", 3), ("2010-01-31", 3)], ["date_col", "offset"] ).withColumn("date_col", col("date_col").cast("date")) # 实现MonthEnd逻辑 df = df.withColumn( "month_end_result", when( col("date_col") == last_day(col("date_col")), last_day(add_months(col("date_col"), col("offset"))) ).otherwise( last_day(add_months(col("date_col"), col("offset") - 1)) ) ) # 输出结果 df.show(truncate=False)
运行结果
输出完全匹配pandas的pd.offsets.MonthEnd(x)逻辑:
+----------+------+----------------+ |date_col |offset|month_end_result| +----------+------+----------------+ |2010-01-15|3 |2010-03-31 | |2010-01-31|3 |2010-04-30 | +----------+------+----------------+
如果偏移量x是固定常量,直接把代码中的col("offset")替换为对应数值即可。
内容的提问来源于stack exchange,提问作者Rens
相关产品推荐
相关产品推荐

