Polars DataFrame按组将元素除以对应组50%分位数实现方案
解决Polars按分组将元素除以组内50%分位数的问题
问题背景
需要将DataFrame中每个分组的Value列元素,除以该分组的50%分位数(中位数),原尝试的代码因维度不匹配无法运行:
df.select(pl.col('Value')) / df.group_by('Group').quantile(.5, 'linear')
示例DataFrame:
import polars as pl df = pl.DataFrame( [ ["A", "A", "A", "A", "B", "B", "B", "B"], [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0], ], schema=["Group", "Value"], )
期望输出每个Value除以对应组中位数后的结果,保留原分组列。
错误原因
原代码中,df.group_by('Group').quantile(.5, 'linear')返回的是每个组一行的DataFrame(共2行),而df.select(pl.col('Value'))是8行的列,两者维度不匹配,无法直接做除法运算。
正确解法
方法1:使用transform结合分组计算
transform会将分组计算的结果广播回原DataFrame的每一行,完美匹配维度:
result = df.with_columns( pl.col('Value') / pl.col('Value').group_by('Group').transform(lambda x: x.quantile(0.5, 'linear')).alias('Value') ) print(result)
方法2:使用窗口函数over更简洁
Polars的窗口函数over可以直接在分组范围内计算统计量,代码更简洁:
result = df.with_columns( pl.col('Value') / pl.col('Value').quantile(0.5, 'linear').over('Group').alias('Value') ) print(result)
输出结果
两种方法都会得到符合预期的结果:
shape: (8, 2) ┌───────┬──────────┐ │ Group ┆ Value │ │ --- ┆ --- │ │ str ┆ f64 │ ╞═══════╪══════════╡ │ A ┆ 0.4 │ │ A ┆ 0.8 │ │ A ┆ 1.2 │ │ A ┆ 1.6 │ │ B ┆ 0.769231 │ │ B ┆ 0.923077 │ │ B ┆ 1.076923 │ │ B ┆ 1.230769 │ └───────┴──────────┘
如果只需要得到计算后的Series,可以单独提取:
value_series = pl.col('Value') / pl.col('Value').quantile(0.5, 'linear').over('Group') # 可以后续用with_columns合并回原DataFrame
内容的提问来源于stack exchange,提问作者AlexanderP
相关产品推荐
相关产品推荐

