You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按指定列索引顺序排序Polars DataFrame(兼容懒加载与立即执行模式)

实现Polars DataFrame按指定索引顺序排序(兼容懒加载/立即执行)

需求说明

需要将DataFrame重新排序,让输出的idx列顺序与原表idx2列的数值顺序完全一致——即输出第N行的idx值等于原表第N行的idx2值。

兼容两种模式的实现方法

核心思路是先建立idx值与行索引的映射,再通过idx2值找到对应的行索引,最后用take()方法重新排列行。这种写法同时支持立即执行(eager)和懒加载(lazy)模式:

import polars as pl

def sort_by_idx_matching(df: pl.DataFrame | pl.LazyFrame) -> pl.DataFrame | pl.LazyFrame:
    # 为原表添加行索引列,用于后续映射
    df_with_row = df.with_row_index("row_idx")
    # 构建idx到行索引的映射表
    lookup_table = df_with_row.select("idx", "row_idx")
    # 通过关联idx2与idx,获取需要提取的行索引顺序
    take_order = df_with_row.join(
        lookup_table,
        left_on="idx2",
        right_on="idx",
        how="left"
    ).select("row_idx_right")
    # 根据行索引顺序重新排列原表
    return df.take(take_order)

验证效果

以题目中的DataFrame为例:

# 原表
df = pl.from_repr("""
shape: (4, 4)
┌─────┬─────┬─────┬──────┐
│ idx ┆ a   ┆ b   ┆ idx2 │
│ --- ┆ --- ┆ --- ┆ ---  │
│ u32 ┆ i64 ┆ i64 ┆ u32  │
╞═════╪═════╪═════╪══════╡
│ 0   ┆ 1   ┆ 4   ┆ 3    │
│ 2   ┆ 1   ┆ 3   ┆ 0    │
│ 3   ┆ 2   ┆ 1   ┆ 2    │
│ 4   ┆ 2   ┆ 2   ┆ 4    │
└─────┴─────┴─────┴──────┘
""")

# 立即执行模式
result_eager = sort_by_idx_matching(df)
print(result_eager)

# 懒加载模式
result_lazy = sort_by_idx_matching(df.lazy()).collect()
print(result_lazy)

执行后会得到期望的输出:

shape: (4, 4)
┌─────┬─────┬─────┬──────┐
│ idx ┆ a   ┆ b   ┆ idx2 │
│ --- ┆ --- ┆ --- ┆ ---  │
│ u32 ┆ i64 ┆ i64 ┆ u32  │
╞═════╪═════╪═════╪══════╡
│ 3   ┆ 2   ┆ 1   ┆ 2    │
│ 0   ┆ 1   ┆ 4   ┆ 3    │
│ 2   ┆ 1   ┆ 3   ┆ 0    │
│ 4   ┆ 2   ┆ 2   ┆ 4    │
└─────┴─────┴─────┴──────┘

原理说明

  1. with_row_index:为每行添加唯一行索引,建立idx值和行位置的关联;
  2. join:通过原表的idx2与映射表的idx匹配,得到每个idx2对应的原表行索引;
  3. take:根据匹配得到的行索引顺序,重新提取原表的行,实现目标排序。

内容的提问来源于stack exchange,提问作者ignoring_gravity

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 18:15:07