You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Polars实现类似Pandas qcut的分位数分箱功能?

用Polars实现类似Pandas qcut的功能

Polars目前没有内置的qcut函数,但可以通过结合现有功能实现相同的分箱效果,以下是两种可行方案:

方案一:基于分位数阈值的分箱

先计算目标分位数的阈值,再用cut函数划分区间,逻辑和Pandas qcut完全对齐:

import polars as pl

# 原始数据
data = pl.Series([11, 1, 2, 2, 3, 4, 5, 1, 2, 3, 4, 5])

# 计算指定分位数的阈值
quantile_points = [0.2, 0.4, 0.6, 0.8]
quantiles = data.quantile(quantile_points)

# 构建分箱边界:首尾添加正负无穷,确保所有数据被包含
bins = [-float("inf")] + quantiles.to_list() + [float("inf")]
labels = ["q1", "q2", "q3", "q4", "q5"]

# 执行分箱
qcut_result = data.cut(bins=bins, labels=labels)
print(qcut_result)

方案二:基于百分位排名的分箱

通过pct_rank计算每个元素的百分位排名,再根据排名区间分配标签,自动处理重复值场景:

import polars as pl

data = pl.Series([11, 1, 2, 2, 3, 4, 5, 1, 2, 3, 4, 5])

qcut_result = data.to_frame("data").with_columns(
    pl.when(pl.col("data").pct_rank() <= 0.2)
    .then(pl.lit("q1"))
    .when(pl.col("data").pct_rank() <= 0.4)
    .then(pl.lit("q2"))
    .when(pl.col("data").pct_rank() <= 0.6)
    .then(pl.lit("q3"))
    .when(pl.col("data").pct_rank() <= 0.8)
    .then(pl.lit("q4"))
    .otherwise(pl.lit("q5"))
    .alias("qcut")
)["qcut"]

print(qcut_result)

补充说明

  • 方案一在数据存在重复值导致分位数点重叠时,会出现空箱,和Pandas qcut未设置duplicates参数的行为一致;若需避免空箱,可手动调整分位数点或合并区间。
  • 方案二会让每个标签对应的元素数量尽可能接近指定比例,适合需要严格按比例划分的场景。

内容的提问来源于stack exchange,提问作者lebesgue

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 05:21:57