You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中筛选满足两个分组条件的完整组数据行

解决Polars中筛选满足双条件的完整分组行问题

你需要筛选同时满足以下两个条件的整个分组所有行:

  • 分组(按item和location)内store的唯一值数量大于1
  • 分组包含特定值(如"store 3")

之前的代码错误在于将行级条件(当前行的store是"store 3")与分组级条件直接拼接,导致仅返回分组内符合行级条件的单行,而非整个分组。

正确思路:将第二个条件也转换为分组级判断——即判断整个分组是否存在目标值,再结合第一个分组条件筛选所有符合的行。


方式一:直接在filter中使用窗口函数

import polars as pl

df = pl.from_repr("""
┌──────┬──────────┬─────────┐
│ item ┆ location ┆ store   │
│ ---  ┆ ---      ┆ ---     │
│ i64  ┆ str      ┆ str     │
╞══════╪══════════╪═════════╡
│ 0    ┆ new york ┆ store 1 │
│ 0    ┆ boston   ┆ store 1 │
│ 1    ┆ boston   ┆ store 1 │
│ 1    ┆ boston   ┆ store 2 │
│ 0    ┆ ohio     ┆ store 1 │
│ 0    ┆ ohio     ┆ store 3 │
└──────┴──────────┴─────────┘
""")

result = df.filter(
    # 条件1:分组内store唯一值数量>1
    (pl.col("store").n_unique() > 1).over("item", "location")
    # 条件2:分组内存在"store 3"(用any()判断分组内是否有符合条件的行)
    & (pl.col("store").is_in(["store 3"]).any().over("item", "location"))
)

print(result)

输出结果:

┌──────┬──────────┬─────────┐
│ item ┆ location ┆ store   │
│ ---  ┆ ---      ┆ ---     │
│ i64  ┆ str      ┆ str     │
╞══════╪══════════╪═════════╡
│ 0    ┆ ohio     ┆ store 1 │
│ 0    ┆ ohio     ┆ store 3 │
└──────┴──────────┴─────────┘

方式二:先添加分组标记列再筛选(可读性更强)

如果觉得窗口函数嵌套在filter里不够直观,可以先通过with_columns生成分组级的标记列,再筛选:

# 添加分组标记列
df_with_flags = df.with_columns(
    has_multiple_stores=(pl.col("store").n_unique() > 1).over("item", "location"),
    contains_store3=(pl.col("store") == "store 3").any().over("item", "location")
)

# 筛选符合条件的行并移除临时标记列
result = df_with_flags.filter(
    pl.col("has_multiple_stores") & pl.col("contains_store3")
).drop(["has_multiple_stores", "contains_store3"])

print(result)

内容的提问来源于stack exchange,提问作者campo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 15:41:23