如何在Polars DataFrame中对指定列执行增量轮换操作?
Polars实现day列增量轮换的惯用方法
可以用两种Polars风格的方式实现需求:
方法一:显式映射字典
先提取排序后的唯一day值,构建映射关系后替换原列:
import polars as pl df = pl.from_repr(""" ┌─────┬───────┐ │ day ┆ value │ │ --- ┆ --- │ │ i64 ┆ i64 │ ╞═════╪═══════╡ │ 1 ┆ 1 │ │ 1 ┆ 2 │ │ 1 ┆ 2 │ │ 3 ┆ 3 │ │ 3 ┆ 5 │ │ 3 ┆ 2 │ │ 5 ┆ 1 │ │ 5 ┆ 2 │ │ 8 ┆ 7 │ │ 8 ┆ 3 │ │ 9 ┆ 5 │ │ 9 ┆ 3 │ │ 9 ┆ 4 │ └─────┴───────┘ """) # 获取排序后的唯一day值列表 sorted_unique_days = df["day"].unique().sort().to_list() # 构建映射:每个day对应下一个更大值,最大值对应None day_mapping = { sorted_unique_days[i]: sorted_unique_days[i+1] if i < len(sorted_unique_days)-1 else None for i in range(len(sorted_unique_days)) } # 替换day列 result = df.with_columns(pl.col("day").replace(day_mapping)) print(result)
方法二:链式表达式操作
利用Polars的链式API,在表达式内完成映射替换,更简洁:
result = df.with_columns( pl.col("day") .map_batches(lambda series: series.replace( dict(zip( series.unique().sort(), series.unique().sort().shift(-1) )) ) ) ) print(result)
两种方法都能得到你要的预期结果,第二种更贴合Polars的函数式编程风格,不需要额外提取变量。
内容的提问来源于stack exchange,提问作者lebesgue
相关产品推荐
相关产品推荐

