You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars 0.18.4中实现自定义频率列表指标计算?

Polars 0.18.4 实现自定义列表计算的方案

问题场景

现有Polars DataFrame如下:

import polars as pl

df = pl.DataFrame({
    "foo": [3, 4, 6],
    "bar_list": [[62, 40, 20], [10, 9, 8, 7], [20, 15, 12, 10, 8, 5]]
})

结构:

shape: (3, 2)
┌─────┬────────────────────────┐
│ foo ┆ bar_list               │
│ --- ┆ ---                    │
│ i64 ┆ list[i64]              │
╞═════╪════════════════════════╡
│ 3   ┆ [62, 40, 20]           │
│ 4   ┆ [10, 9, 8, 7]          │
│ 6   ┆ [20, 15, 12, 10, 8, 5] │
└─────┴────────────────────────┘

其中foo对应bar_list的元素数量,bar_list是降序排列的频率列表。需要实现以下自定义函数的功能:

import math

def get_custom(list_of_frequencies):
    median_freq = math.floor(sum(list_of_frequencies)/2)
    running_total = 0
    counter = 0

    while running_total <= median_freq:
        running_total += list_of_frequencies[counter]
        counter += 1
    
    return round(counter + (1 - (running_total - median_freq)/list_of_frequencies[counter-1]), 2)

在Pandas中可通过df.to_pandas()['bar_list'].apply(get_custom)得到结果:

0    1.98
1    2.78
2    3.00
Name: bar_list, dtype: float64

以下是Polars 0.18.4中的实现方法:


方法一:直接应用自定义函数(简单直观)

Polars支持对列表列直接使用apply方法,传入自定义函数即可,逻辑和Pandas一致:

import polars as pl
import math

def get_custom(list_of_frequencies):
    median_freq = math.floor(sum(list_of_frequencies)/2)
    running_total = 0
    counter = 0

    while running_total <= median_freq:
        running_total += list_of_frequencies[counter]
        counter += 1
    
    return round(counter + (1 - (running_total - median_freq)/list_of_frequencies[counter-1]), 2)

df = pl.DataFrame({
    "foo": [3, 4, 6],
    "bar_list": [[62, 40, 20], [10, 9, 8, 7], [20, 15, 12, 10, 8, 5]]
})

# 添加计算结果列
result_df = df.with_columns(
    pl.col("bar_list").apply(get_custom).alias("custom_result")
)

print(result_df)

输出:

shape: (3, 3)
┌─────┬────────────────────────┬──────────────┐
│ foo ┆ bar_list               ┆ custom_result │
│ --- ┆ ---                    ┆ ---          │
│ i64 ┆ list[i64]              ┆ f64          │
╞═════╪════════════════════════╪══════════════╡
│ 3   ┆ [62, 40, 20]           ┆ 1.98         │
│ 4   ┆ [10, 9, 8, 7]          ┆ 2.78         │
│ 6   ┆ [20, 15, 12, 10, 8, 5] ┆ 3.0          │
└─────┴────────────────────────┴──────────────┘

方法二:Polars表达式实现(高性能推荐)

如果处理大数据量,Python循环的apply性能会受限,推荐用Polars内置的向量化表达式实现逻辑,避免Python层的循环:

import polars as pl

df = pl.DataFrame({
    "foo": [3, 4, 6],
    "bar_list": [[62, 40, 20], [10, 9, 8, 7], [20, 15, 12, 10, 8, 5]]
})

result_df = df.with_columns(
    # 计算总和的一半并向下取整
    median_freq = pl.col("bar_list").list.sum() // 2,
    # 生成累积和列表
    cumulative_sum = pl.col("bar_list").list.cum_sum(),
    # 找到第一个超过median_freq的累积和的索引(转换为从1开始的计数)
    counter = pl.col("cumulative_sum").list.eval(
        pl.arg_where(pl.element() > pl.col("median_freq")).first() + 1
    ),
    # 获取对应位置的频率值
    freq_val = pl.col("bar_list").list.get(pl.col("counter") - 1),
    # 获取目标累积和的值
    target_cum_sum = pl.col("cumulative_sum").list.get(pl.col("counter") - 1),
).with_columns(
    # 计算最终结果并保留两位小数
    custom_result = pl.round(
        pl.col("counter") + (1 - (pl.col("target_cum_sum") - pl.col("median_freq")) / pl.col("freq_val")),
        2
    )
# 清理中间计算列
).drop(["median_freq", "cumulative_sum", "counter", "freq_val", "target_cum_sum"])

print(result_df)

输出结果和方法一完全一致,这种方式利用Polars的内部优化,性能远高于Python循环方案。


内容的提问来源于stack exchange,提问作者benedictine_cumbersome

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 18:33:13