You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars多分类列自定义排序疑问:直接排序与转物理值排序结果不一致?

Polars分类列排序异常:原因与解决方法

这不是Bug,是对Polars分类列的排序逻辑和参数配置理解不到位导致的。

问题根源

Polars的Categorical类型默认按字符串字典序排序,而非你期望的自定义物理存储顺序。你当前的代码只是把目标字符串加入了全局StringCache,但既没给col1、col2指定自定义类别顺序,也没修改分类列的排序规则:

  • 直接调用.sort(["col1", "col2"])时,col2会按字符串字典序(c < h)排列,导致h的行排在c之后,不符合预期。
  • 而.to_physical()会让排序依据分类值的物理索引(也就是你通过pl.Series(["B", "A", "h", "c"]).cast(pl.Categorical)设置的顺序:B=0, A=1, h=2, c=3),因此能得到正确结果。

两种正确解决方法

方法1:转换分类列时指定ordering='physical'

让分类列默认按物理值排序,无需每次排序都调用to_physical():

df = pl.DataFrame({
    "nums": [1, 2, 3, 4],
    "col1": ["A", "B", "B", "A"],
    "col2": ["c", "h", "c", "c"],
})

with pl.StringCache():
    # 注册自定义顺序的类别到缓存
    pl.Series(["B", "A", "h", "c"]).cast(pl.Categorical)
    
    df = df.with_columns(
        pl.col(pl.Utf8).cast(pl.Categorical(ordering="physical"))
    ).sort(["col1", "col2"])
    
    print(df)

方法2:给每个分类列显式指定categories

直接定义每个列的类别顺序,更清晰直观:

df = pl.DataFrame({
    "nums": [1, 2, 3, 4],
    "col1": ["A", "B", "B", "A"],
    "col2": ["c", "h", "c", "c"],
})

df = df.with_columns(
    pl.col("col1").cast(pl.Categorical(categories=["B", "A"], ordering="physical")),
    pl.col("col2").cast(pl.Categorical(categories=["h", "c"], ordering="physical"))
).sort(["col1", "col2"])

print(df)

两种方法都能输出你想要的正确结果:

┌──────┬──────┬──────┐
│ nums ┆ col1 ┆ col2 │
│ ---  ┆ ---  ┆ ---  │
│ i64  ┆ cat  ┆ cat  │
╞══════╪══════╪══════╡
│ 2    ┆ B    ┆ h    │
│ 3    ┆ B    ┆ c    │
│ 1    ┆ A    ┆ c    │
│ 4    ┆ A    ┆ c    │
└──────┴──────┴──────┘

内容的提问来源于stack exchange,提问作者glebcom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 07:24:57