如何在PySpark中按多列分组用均值填充price列缺失值并解决报错
PySpark分组均值填充缺失值并向上取整实现方案
你触发的报错如下:
ValueError: value should be a float, int, long, string, bool or dict
该报错的原因是PySpark的fillna()方法仅支持传入固定值、或者列名到填充值的映射字典,不支持直接传入聚合后的列对象,你要实现的分组级空值填充逻辑需要用窗口函数实现,对应Pandas的transform能力。
完整实现代码
# 导入需要的PySpark函数 from pyspark.sql import functions as F from pyspark.sql.window import Window # 定义分组窗口,按condition、model两列拆分分组 window_spec = Window.partitionBy("condition", "model") # 实现空值填充+向上取整逻辑 cars_new = cars.withColumn( "price", F.ceil( # coalesce作用:第一个参数非空则取原值,为空则取第二个参数的分组均值 F.coalesce( F.col("price"), F.mean(F.col("price")).over(window_spec) ) ) )
逻辑对应说明
和你原有的Pandas实现逻辑完全对齐:
mean(F.col("price")).over(window_spec)对应groupby(['condition', 'model' ])['price'].transform('mean'),为每一行生成分组内的price均值F.coalesce实现fillna的空值替换能力F.ceil完全对应np.ceil的向上取整效果
内容的提问来源于stack exchange,提问作者Foxbat
相关产品推荐
相关产品推荐

