You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Polars高效实现分组后多值用户的行标记?

Polars高效实现分组后每行标记分组内唯一值数量

直接用Polars的**窗口函数(Window Functions)**就能高效满足需求,完全不需要左连接或额外过滤操作,Polars的查询优化器会自动处理广播逻辑,性能拉满。

实现代码

import polars as pl

# 原始LazyFrame
df = pl.LazyFrame({
    'col1': ['a','a','b','b','c','c'],
    'col2': ['undefined','defined','defined','defined','undefined','undefined']
})

# 新增col3列:按col1分组后,计算每组col2的唯一值数量并广播到每行
result_df = df.with_columns(
    pl.col("col2").n_unique().over("col1").alias("col3")
)

# 查看结果
print(result_df.collect())

结果验证

执行后得到的结果完全符合预期:

shape: (6, 3)
┌──────┬───────────┬──────┐
│ col1 ┆ col2      ┆ col3 │
│ ---  ┆ ---       ┆ ---  │
│ str  ┆ str       ┆ u32  │
╞══════╪═══════════╪══════╡
│ a    ┆ undefined ┆ 2    │
│ a    ┆ defined   ┆ 2    │
│ b    ┆ defined   ┆ 1    │
│ b    ┆ defined   ┆ 1    │
│ c    ┆ undefined ┆ 1    │
│ c    ┆ undefined ┆ 1    │
└──────┴───────────┴──────┘

为什么高效?

  • 窗口函数over("col1")会直接在分组维度上计算n_unique(),然后自动将结果广播到组内的每一行,没有额外的中间表或连接操作。
  • 针对LazyFrame,Polars的查询优化器会将窗口操作合并到原查询计划中,避免不必要的数据复制,性能远优于分组后再左连接的方案。

内容的提问来源于stack exchange,提问作者INGl0R1AM0R1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 22:12:14