如何在Polars DataFrame的一次group_by中获取每组首行与末行
在Polars中一次分组同时获取每组首行和末行
你可以通过group_by后调用agg方法,同时指定first()和last()聚合操作,还能给结果字段加前缀/后缀来区分首行和末行的数据,实现一次分组完成需求。
方法一:首行与末行数据整合到同一行
import polars as pl df = pl.DataFrame( { "a": [1, 2, 2, 3, 4, 5], "b": [0.5, 0.5, 4, 10, 14, 13], "c": [True, True, True, False, False, True], "d": ["Apple", "Apple", "Apple", "Banana", "Banana", "Banana"], } ) # 一次分组同时获取首行和末行,用后缀区分字段 result = df.group_by("d", maintain_order=True).agg( pl.all().first().suffix("_first"), pl.all().last().suffix("_last") ) print(result)
输出:
shape: (2, 7) ┌────────┬─────────┬──────────┬───────────┬────────┬─────────┬───────────┐ │ d ┆ a_first ┆ b_first ┆ c_first ┆ a_last ┆ b_last ┆ c_last │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ str ┆ i64 ┆ f64 ┆ bool ┆ i64 ┆ f64 ┆ bool │ ╞════════╪═════════╪══════════╪═══════════╪════════╪═════════╪═══════════╡ │ Apple ┆ 1 ┆ 0.5 ┆ true ┆ 2 ┆ 4.0 ┆ true │ │ Banana ┆ 3 ┆ 10.0 ┆ false ┆ 5 ┆ 13.0 ┆ true │ └────────┴─────────┴──────────┴───────────┴────────┴─────────┴───────────┘
方法二:首行与末行作为单独行展示
如果希望每组的首行和末行是独立的两行,可以用concat合并两次分组结果:
# 分别获取首行和末行后合并 first_rows = df.group_by("d", maintain_order=True).first() last_rows = df.group_by("d", maintain_order=True).last() result_rows = pl.concat([first_rows, last_rows], how="vertical").sort("d") print(result_rows)
输出:
shape: (4, 4) ┌────────┬─────┬──────┬───────┐ │ d ┆ a ┆ b ┆ c │ │ --- ┆ --- ┆ --- ┆ --- │ │ str ┆ i64 ┆ f64 ┆ bool │ ╞════════╪═════╪══════╪═══════╡ │ Apple ┆ 1 ┆ 0.5 ┆ true │ │ Apple ┆ 2 ┆ 4.0 ┆ true │ │ Banana ┆ 3 ┆ 10.0 ┆ false │ │ Banana ┆ 5 ┆ 13.0 ┆ true │ └────────┴─────┴──────┴───────┘
内容的提问来源于stack exchange,提问作者Quinten
相关产品推荐
相关产品推荐

