You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python和Polars按规则对含多列的DataFrame排序

问题:基于父行规则的Polars数据排序需求

原始数据由3组“块”组成,每组包含1条父行(Parent Order ID不为空)和2条子行(Order ID不为空)。需求是根据父行的Direction对子行排序:

  • 若父行Direction为Buy,子行顺序为Sell → Buy
  • 若父行Direction为Sell,子行顺序为Buy → Sell

原始数据代码示例:

import polars as pl
df_original = pl.DataFrame(
    {
        'Direction': ["Buy", "Sell", "Buy", "Sell", "Sell", "Buy"],
        'Order ID': [None, '123_1', '123_0', None, '456_1', '456_0'],
        'Parent Order ID': [123, None, None, 456, None, None],

    }
)

期望的排序后结果:

df_sorted = pl.DataFrame(
    {
        'Direction': ["Buy", "Sell", "Buy", "Sell", "Buy", "Sell"],
        'Order ID': [None, '123_1', '123_0', None, '456_0', '456_1'],
        'Parent Order ID': [123, None, None, 456, None, None],

    }
)

解决方案

核心思路是先标记每个数据块的分组,提取父行的排序规则,再生成对应排序键完成排序:

import polars as pl

df_original = pl.DataFrame(
    {
        'Direction': ["Buy", "Sell", "Buy", "Sell", "Sell", "Buy"],
        'Order ID': [None, '123_1', '123_0', None, '456_1', '456_0'],
        'Parent Order ID': [123, None, None, 456, None, None],
    }
)

df_sorted = (
    df_original
    # 给同一块的行标记相同分组ID:父行出现时初始化,子行继承
    .with_columns(
        group_id=pl.col("Parent Order ID").fill_null(strategy="forward")
    )
    # 将父行的Direction广播到同组所有行,作为排序规则
    .with_columns(
        parent_dir=pl.col("Direction").filter(pl.col("Parent Order ID").is_not_null()).over("group_id")
    )
    # 生成排序键:父行优先排最前,子行按父行规则设置优先级
    .with_columns(
        sort_key=pl.when(pl.col("Parent Order ID").is_not_null())
                  .then(0)
                  .when(pl.col("parent_dir") == "Buy")
                  .then(pl.col("Direction").map({"Sell": 1, "Buy": 2}))
                  .when(pl.col("parent_dir") == "Sell")
                  .then(pl.col("Direction").map({"Buy": 1, "Sell": 2}))
    )
    # 按分组和排序键排序,清理辅助列
    .sort(["group_id", "sort_key"])
    .drop(["group_id", "parent_dir", "sort_key"])
)

print(df_sorted)

逻辑说明

  1. 分组标记:通过forward填充Parent Order ID,让同一数据块的行共享相同group_id,实现分组隔离。
  2. 规则广播:利用Polars的窗口函数over("group_id"),将父行的Direction同步到同组所有子行,统一排序规则。
  3. 排序键生成:给父行设置最低排序键(0)确保排在块首;子行根据父行规则设置对应优先级,满足指定顺序要求。
  4. 最终排序:按分组ID和排序键完成排序后,删除辅助列得到目标结果。

内容的提问来源于stack exchange,提问作者ThoStart

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 22:15:32