You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars使用when/then与选择器时DuplicateError的解决方法

Polars批量处理数值列并保留原列名的解决方案

问题背景

现有如下Polars DataFrame:

import polars as pl
import polars.selectors as cs

df = pl.from_repr("""
┌─────┬─────┬─────┐
│ j   ┆ k   ┆ l   │
│ --- ┆ --- ┆ --- │
│ i64 ┆ i64 ┆ i64 │
╞═════╪═════╪═════╡
│ 71  ┆ 79  ┆ 67  │
│ 26  ┆ 42  ┆ 55  │
│ 12  ┆ 43  ┆ 85  │
│ 92  ┆ 96  ┆ 14  │
│ 95  ┆ 26  ┆ 62  │
│ 75  ┆ 14  ┆ 56  │
│ 61  ┆ 41  ┆ 75  │
│ 74  ┆ 97  ┆ 70  │
│ 73  ┆ 32  ┆ 10  │
│ 66  ┆ 98  ┆ 40  │
└─────┴─────┴─────┘
""")

尝试用cs.numeric()批量对数值列应用when/then/otherwise逻辑时:

df.select(
    pl.when(cs.numeric() < 50)
      .then(1)
      .otherwise(2)
)

会触发DuplicateError: the name 'literal' is duplicate——原因是所有处理后的列默认都命名为literal,导致列名重复。需要找到更简洁的方法,替代手动遍历列并指定别名的代码,同时保留原列名:

df.select(
    pl.when(pl.col(c) < 50)
      .then(1)
      .otherwise(2)
      .alias(c)
    for c in df.columns
)

解决方案

使用Polars选择器的map方法,对选中的每一列批量应用逻辑,自动保留原列名:

df.select(
    cs.numeric().map(lambda col: pl.when(col < 50).then(1).otherwise(2))
)

结果验证

执行后得到的结果与手动遍历的效果完全一致:

shape: (10, 3)
┌─────┬─────┬─────┐
│ j   ┆ k   ┆ l   │
│ --- ┆ --- ┆ --- │
│ i32 ┆ i32 ┆ i32 │
╞═════╪═════╪═════╡
│ 2   ┆ 2   ┆ 2   │
│ 1   ┆ 1   ┆ 2   │
│ 1   ┆ 1   ┆ 2   │
│ 2   ┆ 2   ┆ 1   │
│ 2   ┆ 1   ┆ 2   │
│ 2   ┆ 1   ┆ 2   │
│ 2   ┆ 1   ┆ 2   │
│ 2   ┆ 2   ┆ 2   │
│ 2   ┆ 1   ┆ 1   │
│ 2   ┆ 2   ┆ 1   │
└─────┴─────┴─────┘

内容的提问来源于stack exchange,提问作者levant pied

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 04:52:43