如何按指定列索引顺序排序Polars DataFrame(兼容懒加载与立即执行模式)
实现Polars DataFrame按指定索引顺序排序(兼容懒加载/立即执行)
需求说明
需要将DataFrame重新排序,让输出的idx列顺序与原表idx2列的数值顺序完全一致——即输出第N行的idx值等于原表第N行的idx2值。
兼容两种模式的实现方法
核心思路是先建立idx值与行索引的映射,再通过idx2值找到对应的行索引,最后用take()方法重新排列行。这种写法同时支持立即执行(eager)和懒加载(lazy)模式:
import polars as pl def sort_by_idx_matching(df: pl.DataFrame | pl.LazyFrame) -> pl.DataFrame | pl.LazyFrame: # 为原表添加行索引列,用于后续映射 df_with_row = df.with_row_index("row_idx") # 构建idx到行索引的映射表 lookup_table = df_with_row.select("idx", "row_idx") # 通过关联idx2与idx,获取需要提取的行索引顺序 take_order = df_with_row.join( lookup_table, left_on="idx2", right_on="idx", how="left" ).select("row_idx_right") # 根据行索引顺序重新排列原表 return df.take(take_order)
验证效果
以题目中的DataFrame为例:
# 原表 df = pl.from_repr(""" shape: (4, 4) ┌─────┬─────┬─────┬──────┐ │ idx ┆ a ┆ b ┆ idx2 │ │ --- ┆ --- ┆ --- ┆ --- │ │ u32 ┆ i64 ┆ i64 ┆ u32 │ ╞═════╪═════╪═════╪══════╡ │ 0 ┆ 1 ┆ 4 ┆ 3 │ │ 2 ┆ 1 ┆ 3 ┆ 0 │ │ 3 ┆ 2 ┆ 1 ┆ 2 │ │ 4 ┆ 2 ┆ 2 ┆ 4 │ └─────┴─────┴─────┴──────┘ """) # 立即执行模式 result_eager = sort_by_idx_matching(df) print(result_eager) # 懒加载模式 result_lazy = sort_by_idx_matching(df.lazy()).collect() print(result_lazy)
执行后会得到期望的输出:
shape: (4, 4) ┌─────┬─────┬─────┬──────┐ │ idx ┆ a ┆ b ┆ idx2 │ │ --- ┆ --- ┆ --- ┆ --- │ │ u32 ┆ i64 ┆ i64 ┆ u32 │ ╞═════╪═════╪═════╪══════╡ │ 3 ┆ 2 ┆ 1 ┆ 2 │ │ 0 ┆ 1 ┆ 4 ┆ 3 │ │ 2 ┆ 1 ┆ 3 ┆ 0 │ │ 4 ┆ 2 ┆ 2 ┆ 4 │ └─────┴─────┴─────┴──────┘
原理说明
with_row_index:为每行添加唯一行索引,建立idx值和行位置的关联;join:通过原表的idx2与映射表的idx匹配,得到每个idx2对应的原表行索引;take:根据匹配得到的行索引顺序,重新提取原表的行,实现目标排序。
内容的提问来源于stack exchange,提问作者ignoring_gravity
相关产品推荐
相关产品推荐

