You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars中Pivot操作后高效重命名DataFrame列的方法

Polars透视后列名重命名的通用解决方案

Polars执行pivot操作时,默认会将值列名与透视列(on参数指定)的取值用下划线拼接作为新列名,例如monthly_qty_product_a。若需要将列名格式调整为product_a:monthly_qty(透视列值在前,值列名在后,用冒号分隔),且要兼容值列名包含多个下划线的场景(如monthly_growth_pct),可以用以下两种更通用的实现方式:

方法一:透视时直接指定列名生成规则(最优)

利用Polars的pivot方法提供的names_generator参数,直接定义列名的生成逻辑,无需事后重命名,效率更高。

代码示例

import polars as pl

# 构造测试数据,包含带多个下划线的值列
data = pl.DataFrame({
    "month": ["Jan", "Jan", "Feb", "Feb", "Mar", "Mar"],
    "type": ["product_a", "product_b"] * 3,
    "monthly_qty": [10, 20] * 3,
    "monthly_amt": [5., 8.] * 3,
    "monthly_growth_pct": [0.1, 0.2] * 3
})

# 透视时指定列名生成规则:透视列值:值列名
pivoted_data = data.pivot(
    on="type",
    index="month",
    names_generator=lambda value_col, pivot_value: f"{pivot_value}:{value_col}"
)

print(pivoted_data)

输出效果

shape: (3, 7)
┌───────┬───────────────────────┬───────────────────────┬───────────────────────┬───────────────────────┬───────────────────────────┬───────────────────────────┐
│ month ┆ product_a:monthly_qty ┆ product_b:monthly_qty ┆ product_a:monthly_amt ┆ product_b:monthly_amt ┆ product_a:monthly_growth_pct ┆ product_b:monthly_growth_pct │
│ ---   ┆ ---                   ┆ ---                   ┆ ---                   ┆ ---                   ┆ ---                         ┆ ---                         │
│ str   ┆ i64                   ┆ i64                   ┆ f64                   ┆ f64                   ┆ f64                         ┆ f64                         │
╞═══════╪═══════════════════════╪═══════════════════════╪═══════════════════════╪═══════════════════════╪═════════════════════════════╪═════════════════════════════╡
│ Jan   ┆ 10                    ┆ 20                    ┆ 5.0                   ┆ 8.0                   ┆ 0.1                         ┆ 0.2                         │
│ Feb   ┆ 10                    ┆ 20                    ┆ 5.0                   ┆ 8.0                   ┆ 0.1                         ┆ 0.2                         │
│ Mar   ┆ 10                    ┆ 20                    ┆ 5.0                   ┆ 8.0                   ┆ 0.1                         ┆ 0.2                         │
└───────┴───────────────────────┴───────────────────────┴───────────────────────┴───────────────────────┴───────────────────────────┴───────────────────────────┘

方法二:对已透视的DataFrame批量重命名

如果已经完成透视操作,可通过拆分最后一个下划线的方式实现通用重命名(因为Polars默认拼接规则是值列名_透视列值,最后一个下划线是两者的分界)。

代码示例

# 假设已有默认透视后的DataFrame
default_pivoted = data.pivot(on="type", index="month")

# 批量生成新列名映射
new_col_mapping = {}
for col in default_pivoted.columns:
    if col == "month":
        new_col_mapping[col] = col
    else:
        # 找到最后一个下划线的位置,拆分出两部分
        split_pos = col.rfind("_")
        value_col_name = col[:split_pos]
        pivot_value = col[split_pos+1:]
        new_col_mapping[col] = f"{pivot_value}:{value_col_name}"

# 执行重命名
renamed_data = default_pivoted.rename(new_col_mapping)
print(renamed_data)

该方法同样能正确处理值列名含多个下划线的场景,拆分逻辑不受值列名下划线数量影响。


内容的提问来源于stack exchange,提问作者JASMINE LIAW

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 20:20:29