如何同时对两个Polars DataFrame执行分组操作?
解决方案
要实现两个DataFrame按相同键分组后保留各自组内数据的效果,正确的做法是分别对两个DataFrame分组聚合,再通过分组键做内连接,而非使用with_context。这种方式符合Polars官方API规范,且能直接迁移到Rust实现。
代码实现
import polars as pl df = pl.DataFrame({ "grpbyKey": [1, 1, 1, 2, 2, 2], "val": ["One"] * 3 + ["Two"] * 3 }) df2 = pl.DataFrame({ "grpbyKey": [1, 1, 2, 2, 2, 3], "val2": ["One"] * 2 + ["Two"] * 3 + ["Three"] }) # 分别对两个DataFrame分组聚合 df_grouped = df.lazy().group_by("grpbyKey").agg(pl.col("val").alias("val")).collect() df2_grouped = df2.lazy().group_by("grpbyKey").agg(pl.col("val2").alias("val2")).collect() # 内连接得到目标结果 result = df_grouped.join(df2_grouped, on="grpbyKey", how="inner") print(result)
执行结果
shape: (2, 3) ┌──────────┬───────────────────────┬───────────────────────┐ │ grpbyKey ┆ val ┆ val2 │ │ --- ┆ --- ┆ --- │ │ i64 ┆ list[str] ┆ list[str] │ ╞══════════╪═══════════════════════╪═══════════════════════╡ │ 1 ┆ ["One", "One", "One"] ┆ ["One", "One"] │ │ 2 ┆ ["Two", "Two", "Two"] ┆ ["Two", "Two", "Two"] │ └──────────┴───────────────────────┴───────────────────────┘
后续自定义函数扩展
如果需要在分组内执行自定义逻辑,可以使用map_groups分别处理两个DataFrame的分组,再连接结果。示例如下:
# 自定义分组处理函数 def process_df_group(group): return group.select( "grpbyKey", pl.col("val").list.sort().alias("val") # 示例自定义操作:排序列表 ) def process_df2_group(group): return group.select( "grpbyKey", pl.col("val2").list.unique().alias("val2") # 示例自定义操作:去重 ) # 分组处理后连接 df_processed = df.lazy().group_by("grpbyKey").map_groups(process_df_group).collect() df2_processed = df2.lazy().group_by("grpbyKey").map_groups(process_df2_group).collect() result = df_processed.join(df2_processed, on="grpbyKey", how="inner") print(result)
原方法问题说明
你之前使用with_context的方式,本质是将df2作为查询上下文添加到df的查询中,分组时会把df2的所有数据纳入聚合范围,导致val2列包含了非当前组的数据(比如grpbyKey=1的val2包含了df2中grpbyKey=2的"Two"),这不符合你期望的"同时分组"逻辑。
内容的提问来源于stack exchange,提问作者The Unfun Cat
相关产品推荐
相关产品推荐

