You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Polars中实现多级索引表头扁平化(支持自定义分隔符)的原生方案

在Polars中实现多级索引表头扁平化(支持自定义分隔符)的原生方案

当然有原生的Polars方案可以直接搞定这个需求!不用来回切换Pandas,大数据集下性能提升会非常明显——毕竟Polars在内存效率和处理速度上本来就比Pandas有优势。下面我就一步步带你实现:


一、直接用Polars处理带多级表头的数据源(比如Excel)

你的场景是模拟Excel的多级表头输入,那我们可以直接用Polars读取并处理,全程不需要碰Pandas:

1. 读取多级表头文件

假设你的Excel有两行表头(第一层级是["A", "A", "B", "B"],第二层级是["one", "two", "one", "two"]),用Polars读取时可以指定表头行数,再自定义分隔符扁平化列名:

import polars as pl

# 实际场景直接替换成你的Excel路径
df = pl.read_excel("your_multiheader_file.xlsx", header_rows=[0, 1])

# 默认Polars会用下划线连接多级表头(比如"A_one"),如果要改成你需要的"-"分隔,直接重命名列
sep = "-"
df_flat = df.rename({col: col.replace("_", sep) for col in df.columns})
print(df_flat)

输出结果就是你要的扁平化表头Polars DataFrame:

shape: (2, 4)
┌───────┬───────┬───────┬───────┐
│ A-one ┆ A-two ┆ B-one ┆ B-two │
│ ---   ┆ ---   ┆ ---   ┆ ---   │
│ i64   ┆ i64   ┆ i64   ┆ i64   │
╞═══════╪═══════╪═══════╪═══════╡
│ 1     ┆ 2     ┆ 3     ┆ 4     │
│ 5     ┆ 6     ┆ 7     ┆ 8     │
└───────┴───────┴───────┴───────┘

2. 手动构造多级表头数据并扁平化(如果不是从文件读取)

如果是像你那样手动构造类似的多级表头数据,也可以用Polars直接处理:

import polars as pl
from io import StringIO

# 模拟带多级表头的原始数据
raw_content = """
,A,A,B,B
,one,two,one,two
0,1,2,3,4
1,5,6,7,8
"""

# 先读取所有行(不识别表头)
temp_df = pl.read_csv(StringIO(raw_content), has_header=False)

# 提取前两行作为表头层级
level1 = temp_df.row(0)[1:]  # 跳过第一列的索引列
level2 = temp_df.row(1)[1:]

# 自定义分隔符构造扁平化列名
sep = "-"
flat_columns = [f"{l1}{sep}{l2}" for l1, l2 in zip(level1, level2)]

# 提取数据行并重命名列
df_flat = temp_df.slice(2)  # 从第三行开始是数据
df_flat = df_flat.rename(dict(zip(temp_df.columns[1:], flat_columns))).drop("column_0")

print(df_flat)

二、从扁平化Polars DataFrame还原多级表头(用于导出Excel)

如果你之后需要还原成多级表头导出Excel,也可以用Polars处理后转Pandas(毕竟Pandas在导出多级表头Excel上更成熟):

import pandas as pd

def polars_flat_to_multiindex(df: pl.DataFrame, sep: str = "-") -> pd.DataFrame:
    # 拆分扁平化列名为多级元组
    column_tuples = [col.split(sep, 1) if sep in col else (col, "") for col in df.columns]
    # 转Pandas并设置多级列索引
    pandas_multi_df = df.to_pandas()
    pandas_multi_df.columns = pd.MultiIndex.from_tuples(column_tuples)
    return pandas_multi_df

# 还原示例
pandas_multi_df = polars_flat_to_multiindex(df_flat)
print(pandas_multi_df)

输出就是你要的多级索引DataFrame:

A       B    
  one two one two
0   1   2   3   4
1   5   6   7   8

为什么推荐原生Polars方案?

和你之前来回切换Pandas的方法比,原生Polars处理有两个核心优势:

  1. 性能提升:大数据集下,Polars的内存占用更低,处理速度更快,避免了Pandas和Polars之间的数据转换开销。
  2. 代码简洁:全程用Polars的API处理,不需要额外写Pandas的转换函数,逻辑更连贯。

备注:内容来源于stack exchange,提问作者Reveur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 16:58:07