You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark DataFrame中pandas add_prefix与reset_index的替代实现方法

PySpark对应实现方案
  • reset_index()等效逻辑:Pandas中调用pivot时设置的index="msisdn"会把msisdn转为行索引,因此需要reset_index()恢复为普通列。但PySpark的groupBy("msisdn").pivot(...)逻辑中,分组键msisdn默认就会作为普通列保留在结果集中,不需要额外做重置索引操作。
  • add_prefix('arpu_sum_l')等效逻辑:可以通过遍历结果列名,对非分组键的列批量重命名实现前缀添加,依赖PySpark的select方法批量设置别名即可。

完整实现代码如下:

# 提前导入依赖(如果未导入的话)
import pyspark.sql.functions as F

# 原始pivot逻辑
cdr = datamonthly.groupBy("msisdn").pivot("last_x_month").sum("arpu_sum")
# 批量添加前缀,等价于pandas的add_prefix方法
cdr = cdr.select(
    "msisdn",
    *[F.col(col).alias(f"arpu_sum_l{col}") for col in cdr.columns if col != "msisdn"]
)

内容的提问来源于stack exchange,提问作者Nabih Bawazir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 19:27:03