You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Polars中.hist()返回的分类列区间值?

处理Polars直方图结果的category列筛选

先整理你给出的代码及输出:

import polars as pl
import numpy as np

df = pl.DataFrame(dict(j=np.random.randint(10, 99, 20)))
print(df)
# 输出:
# shape: (20, 1)
# ┌─────┐
# │ j   │
# │ --- │
# │ i64 │
# ╞═════╡
# │ 47  │
# │ 22  │
# │ 82  │
# │ 19  │
# │ …   │
# │ 28  │
# │ 94  │
# │ 21  │
# │ 38  │
# └─────┘

hist_df = df.get_column('j').hist([10, 20, 30, 50])
print(hist_df)
# 输出:
# shape: (5, 3)
# ┌─────────────┬──────────────┬─────────┐
# │ break_point ┆ category     ┆ j_count │
# │ ---         ┆ ---          ┆ ---     │
# │ f64         ┆ cat          ┆ u32     │
# ╞═════════════╪══════════════╪═════════╡
# │ 10.0        ┆ (-inf, 10.0] ┆ 0       │
# │ 20.0        ┆ (10.0, 20.0] ┆ 4       │
# │ 30.0        ┆ (20.0, 30.0] ┆ 5       │
# │ 50.0        ┆ (30.0, 50.0] ┆ 3       │
# │ inf         ┆ (50.0, inf]  ┆ 8       │
# └─────────────┴──────────────┴─────────┘

以下是针对需求的具体操作:

1. 筛选包含-inf的分类

直接对category列做字符串匹配即可,Polars支持对分类列直接执行字符串操作:

# 筛选含-inf的行
filtered_inf = hist_df.filter(pl.col("category").str.contains("-inf"))
print(filtered_inf)

执行后输出:

shape: (1, 3)
┌─────────────┬──────────────┬─────────┐
│ break_point ┆ category     ┆ j_count │
│ ---         ┆ ---          ┆ ---     │
│ f64         ┆ cat          ┆ u32     │
╞═════════════╪══════════════╪═════════╡
│ 10.0        ┆ (-inf, 10.0] ┆ 0       │
└─────────────┴──────────────┴─────────┘

2. 筛选上限在10.0到30.0之间的分类

推荐两种高效实现方式:

方式一:利用break_point列(更简洁)

观察结果可知,break_point就是对应区间的上限值(最后一行的inf除外),直接筛选该列的范围即可:

# 筛选上限在10到30之间的分类
filtered_upper = hist_df.filter(pl.col("break_point").is_between(10, 30, closed="right"))
print(filtered_upper)

执行后输出:

shape: (2, 3)
┌─────────────┬──────────────┬─────────┐
│ break_point ┆ category     ┆ j_count │
│ ---         ┆ ---          ┆ ---     │
│ f64         ┆ cat          ┆ u32     │
╞═════════════╪══════════════╪═════════╡
│ 20.0        ┆ (10.0, 20.0] ┆ 4       │
│ 30.0        ┆ (20.0, 30.0] ┆ 5       │
└─────────────┴──────────────┴─────────┘

方式二:解析category字符串提取上限(适合特殊场景)

如果必须从category字符串中提取上限再筛选,可使用正则匹配:

# 提取区间上限并筛选
filtered_upper_str = hist_df.with_columns(
    # 用正则提取区间的上限值
    upper_bound=pl.col("category").str.extract(r",\s*([\d.]+|\w+)\]").cast(pl.Float64)
).filter(pl.col("upper_bound").is_between(10, 30))

print(filtered_upper_str)

该方法会新增一列upper_bound存储提取到的上限值,筛选结果与方式一一致。


内容的提问来源于stack exchange,提问作者levant pied

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 06:35:11