You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark DataFrame条件过滤与分层比例插补需求

PySpark DataFrame 零值插补需求

现有DataFrame结构与数据

给定PySpark DataFrame df,结构及数据如下:

IDTotal_CountABCDGroupNameChain
15600001AppleFruits1
26500001AppleFruits1
372003001BananaFruits1
48000001StrawberryFruits1
5142585814121AppleFruits1
61306350981AppleFruits1
7145744417101AppleFruits1
81195448891AppleFruits1
11161716316111BananaFruits1
1212454431981BananaFruits1

插补逻辑

需要对A、B、C、D列中存在零值的行(ID为1、2、3、4的行)进行插补,规则如下:

  1. 优先采用Group×Name层级的比例平均值(该层级有非零数据时),其次采用Group×Chain层级的比例平均值,最后采用Group层级的比例平均值。
    • 以插补ID1、2为例:筛选Group=1且Name=Apple的非零行,计算每行的比例A/Total_Count、B/Total_Count、C/Total_Count、D/Total_Count,得到比例表:
      A_PROPB_PROPC_PROPD_PROP
      0.4084510.4084507040.0985920.084507042
      0.4846150.3846153850.0692310.061538462
      0.5103450.3034482760.1172410.068965517
      0.4537820.4033613450.0672270.075630252
  2. 取上述比例的平均值,将平均值乘以当前行的Total_Count,得到插补后的A、B、C、D值(结果取整)。

预期输出DataFrame df2

IDTotal_CountA_prop_avgB_prop_avgC_prop_avgD_prop_avgABCD
1560.464298110.374968930.088072650.07266032262154
2650.4642981070.3749689270.0880726470.072660318302465
3720.438238830.3690392710.1263023440.066419555322795
4800.4556116810.3729923750.100815880.070580064363086

内容的提问来源于stack exchange,提问作者Scope

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 01:50:24