如何在Polars 0.18.4中实现自定义频率列表指标计算?
Polars 0.18.4 实现自定义列表计算的方案
问题场景
现有Polars DataFrame如下:
import polars as pl df = pl.DataFrame({ "foo": [3, 4, 6], "bar_list": [[62, 40, 20], [10, 9, 8, 7], [20, 15, 12, 10, 8, 5]] })
结构:
shape: (3, 2) ┌─────┬────────────────────────┐ │ foo ┆ bar_list │ │ --- ┆ --- │ │ i64 ┆ list[i64] │ ╞═════╪════════════════════════╡ │ 3 ┆ [62, 40, 20] │ │ 4 ┆ [10, 9, 8, 7] │ │ 6 ┆ [20, 15, 12, 10, 8, 5] │ └─────┴────────────────────────┘
其中foo对应bar_list的元素数量,bar_list是降序排列的频率列表。需要实现以下自定义函数的功能:
import math def get_custom(list_of_frequencies): median_freq = math.floor(sum(list_of_frequencies)/2) running_total = 0 counter = 0 while running_total <= median_freq: running_total += list_of_frequencies[counter] counter += 1 return round(counter + (1 - (running_total - median_freq)/list_of_frequencies[counter-1]), 2)
在Pandas中可通过df.to_pandas()['bar_list'].apply(get_custom)得到结果:
0 1.98 1 2.78 2 3.00 Name: bar_list, dtype: float64
以下是Polars 0.18.4中的实现方法:
方法一:直接应用自定义函数(简单直观)
Polars支持对列表列直接使用apply方法,传入自定义函数即可,逻辑和Pandas一致:
import polars as pl import math def get_custom(list_of_frequencies): median_freq = math.floor(sum(list_of_frequencies)/2) running_total = 0 counter = 0 while running_total <= median_freq: running_total += list_of_frequencies[counter] counter += 1 return round(counter + (1 - (running_total - median_freq)/list_of_frequencies[counter-1]), 2) df = pl.DataFrame({ "foo": [3, 4, 6], "bar_list": [[62, 40, 20], [10, 9, 8, 7], [20, 15, 12, 10, 8, 5]] }) # 添加计算结果列 result_df = df.with_columns( pl.col("bar_list").apply(get_custom).alias("custom_result") ) print(result_df)
输出:
shape: (3, 3) ┌─────┬────────────────────────┬──────────────┐ │ foo ┆ bar_list ┆ custom_result │ │ --- ┆ --- ┆ --- │ │ i64 ┆ list[i64] ┆ f64 │ ╞═════╪════════════════════════╪══════════════╡ │ 3 ┆ [62, 40, 20] ┆ 1.98 │ │ 4 ┆ [10, 9, 8, 7] ┆ 2.78 │ │ 6 ┆ [20, 15, 12, 10, 8, 5] ┆ 3.0 │ └─────┴────────────────────────┴──────────────┘
方法二:Polars表达式实现(高性能推荐)
如果处理大数据量,Python循环的apply性能会受限,推荐用Polars内置的向量化表达式实现逻辑,避免Python层的循环:
import polars as pl df = pl.DataFrame({ "foo": [3, 4, 6], "bar_list": [[62, 40, 20], [10, 9, 8, 7], [20, 15, 12, 10, 8, 5]] }) result_df = df.with_columns( # 计算总和的一半并向下取整 median_freq = pl.col("bar_list").list.sum() // 2, # 生成累积和列表 cumulative_sum = pl.col("bar_list").list.cum_sum(), # 找到第一个超过median_freq的累积和的索引(转换为从1开始的计数) counter = pl.col("cumulative_sum").list.eval( pl.arg_where(pl.element() > pl.col("median_freq")).first() + 1 ), # 获取对应位置的频率值 freq_val = pl.col("bar_list").list.get(pl.col("counter") - 1), # 获取目标累积和的值 target_cum_sum = pl.col("cumulative_sum").list.get(pl.col("counter") - 1), ).with_columns( # 计算最终结果并保留两位小数 custom_result = pl.round( pl.col("counter") + (1 - (pl.col("target_cum_sum") - pl.col("median_freq")) / pl.col("freq_val")), 2 ) # 清理中间计算列 ).drop(["median_freq", "cumulative_sum", "counter", "freq_val", "target_cum_sum"]) print(result_df)
输出结果和方法一完全一致,这种方式利用Polars的内部优化,性能远高于Python循环方案。
内容的提问来源于stack exchange,提问作者benedictine_cumbersome
相关产品推荐
相关产品推荐

